Skip to content
View taipei49314's full-sized avatar

Block or report taipei49314

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
taipei49314/README.md

Nelson — evidence-first agent systems

I build local-first, deterministic, auditable tools for AI agents and developer workflows.

我在台灣打造本機優先、可重播、可稽核的 AI agent 與開發工具。

Models may propose. Verifiers decide. Missing evidence stays UNKNOWN / INCOMPLETE.

Explore the full Nelson Stack →

Admission then measure — unsigned walkaround, refused advance

Now

The active line is admit, then measure:

  • walkaround asks whether the session entered a frozen contract. Done without entry is BYPASSED. Receipts are unsigned; no VERIFIED.
  • charterlock asks whether that journey was allowed to be the exam. Same key writing and sitting it is CHARTER_COLLAPSED. Two MAC keys do not prove two people.
  • trust-meter scores a checkout from evidence.
  • phaseledger refuses to advance a phase without a fresh measurer PASS.
  • nullbench pre-registers decisions and scores them against chance — never backfill.

No GitHub Release for walkaround / charterlock / trust-meter / phaseledger. Lab publication remains a measured pre-alpha with explicit NOT_RUN / BLOCKED rows.

The flagship is that audit spine, not a demo. First CI-backed public subject: cell-shift (CELL//SHIFT). Frozen journey and external receipts live in that repo. The chamber's own tests are not the spine verdict.

The audit loop

Audit loop with honest gaps

Stage Project Question
Admit walkaround Did this session enter a frozen task contract?
Charter charterlock Was the journey allowed to count as an exam?
Measure trust-meter What does this checkout actually score from local evidence?
Gate phaseledger Can this phase advance without a fresh deterministic PASS?
Detect greenwash Did an agent make CI green by weakening verification?
Baseline nullbench Did the decision beat chance, with the claim pre-registered?
Re-run RepoPassport Did the declared journey work within its capabilities and clean up after itself?
Reproduce stateweaver Can an independent verifier replay the finding against a clean state?
Forecast tomorrowci · tomorrowci-lab When will dependency or runtime drift invalidate today’s evidence?

Supporting evaluation surfaces:

  • unasked — evidence-gated repo investigation; non-certifying
  • smallestlie — authorized adversarial harness; smallest accepted lie
  • null-city — deterministic crisis-response sandbox
  • NormShift — evidence-backed standards diffs
  • branchback — decision replay lab; belief-at-the-time vs knowledge-now
  • constraint-deck — session-first authorial constraint deck; measure first

Measured public surfaces

Project Current public status
cell-shift Public subject, not the flagship. Deterministic 3D tissue chamber (CELL//SHIFT). Maintenance-only. Host CI is green; spine verify is not claimed. Not biology.
greenwash v0.1.41 — deterministic diff-level verification-tampering detector. At the 2026-08-15 reconciliation, main was one README-only Action-pin commit ahead of that tag.
nullbench v0.7.0 — pre-register decisions; score against chance; never backfill
tomorrowci-lab v0.2.0-alpha.1 — project-operated prerelease; CANDIDATE_ONLY_NOT_RELEASE_AUTHORIZED. GitHub “Latest” still points at rejected v0.1.0-grok-session
unasked v0.4.0 — authenticated trust plane; public result remains M0_NOT_DEMONSTRATED
null-city v0.1.0-alpha.1 — playable deterministic agent-evaluation sandbox (prerelease; no GitHub “Latest”)
receiptradar v0.1.0-cli.34 — four-platform CLI release with packaged checksums
md-brain v0.2.0 — public prototype for model-independent Markdown continuity (prerelease; no GitHub “Latest”)
aurora v0.1.47 — alpha evidence engine; observations, not investment claims
github-radar v0.1.0 — reproducible stdlib-only alpha with bounded coverage claims (prerelease; no GitHub “Latest”)
nelson-release-studio v1.0.0 — verified Windows-first local release workbench
FutureShow-pet v0.1.0 — personal Windows alpha; Loop 10 and long-soak automation remain open (prerelease; no GitHub “Latest”)
tw-stock-lab v0.2.0 — local TW stock research lab; paper simulation only, not investment advice
branchback v2.0.0 — local-first decision replay laboratory

tomorrowci still has GitHub “Latest” pointed at v0.1.0-grok-session. That tag is a rejected historical candidate, not an acceptance claim. Current lab publication lives in tomorrowci-lab as the v0.2.0-alpha.1 candidate-only prerelease.

Active qualification tracks

Project Honest boundary
walkaround No release. M4 local kernel; receipts unsigned; ADMITTED is not verified work
charterlock No release. independence_claim is always not_claimed; two MAC keys do not prove two people
trust-meter No release. Local scorer with batch/compare/API surfaces; self-audit only
phaseledger No release. Phase advance requires a fresh measurer PASS; reclaim invalidates later phases
tomorrowci-lab Newest prerelease is v0.2.0-alpha.1 and remains candidate-only. macOS / Windows clean-machine and independent authorization remain BLOCKED
RepoPassport Working v1alpha1 vertical slice; 37-row acceptance registry is machine-checked; observer coverage remains incomplete, so healthy runs stay inconclusive
stateweaver Source-only pre-alpha; M6–M8 implementation gates exist; trusted Reality proof does not
NormShift M0 implemented; production and release remain blocked pending external audit
smallestlie Authorized adversarial harness; no release
constraint-deck Public source; measure-first voice contract; no release yet
editorial-doll-engineering-preview Public M0–M3 engineering preview of a deterministic styling engine; no release yet
universe-explorer Public epistemically honest science knowledge system; no release yet
vibe-oracle Explicitly not evidence — vibe theater that admits the theater
why-ledger Justified sovereign decisions notebook (WJSD); documentation-first; no release yet

Local-first tools

Project What it does
nelsoncode-ide Off-mainline optional personal AI coding preview; external security audit remains NO-GO for untrusted use
receiptradar Receipt-to-ledger CLI with no cloud account
md-brain Model-independent continuity runtime for Markdown memory
github-radar GitHub research with measured uncertainty and zero runtime dependencies
aurora Finds unnamed industries from evidence, with no LLM at runtime
music-lab Deterministic local music toolkit; analysis first, no cloud account
nelson-release-studio Windows-first music, asset, lyric-video, and release-package workbench
FutureShow-pet Personal Windows desktop pet with Taiwan and GitHub information loops
tw-stock-lab Local TW market research desk (charts, sector map, signals + Ollama); paper only
branchback Preserve belief-at-the-time vs knowledge-now for decision replay

Engineering rules

  1. Deterministic first. The same evidence should produce the same verdict.
  2. Evidence over vibe. Claims point to tests, diffs, logs, artifacts, or an explicit insufficient-data result.
  3. Fail closed. Missing observation is not a pass; capability violations outrank functional success.
  4. Local first. Prefer loopback services, offline-capable CLIs, and local runtimes over mandatory cloud accounts.
  5. Keep the failures. Negative controls, blocked gates, and NO-GO verdicts remain visible instead of being rewritten as success.

Stack

Python · Rust · Go · TypeScript · FastAPI · React · Electron · SQLite · Ollama

Older market, persona, creative-production, and agent-console experiments are kept private or archived when they stop being the active line. Market-related projects are paper-only research simulations, never broker or investment systems.

Last portfolio reconciliation: 2026-08-17. Aligned with nelson-stack.

Pinned Loading

  1. greenwash greenwash Public

    Catch the agent that made CI green by weakening the tests. Deterministic, zero-LLM, zero-network, sub-second — reads the diff, not the code state. Public bypass list included.

    Python

  2. stateweaver stateweaver Public

    Reality-tethered, state-first security research for authorized synthetic labs.

    Python

  3. universe-explorer universe-explorer Public

    Honestly separating what we know from what we don't — an epistemically honest science knowledge system. 誠實區分已知與未知的科學知識系統。

    Python 1

  4. editorial-doll-engineering-preview editorial-doll-engineering-preview Public

    Public engineering preview of a deterministic editorial styling engine (M0-M3).

    TypeScript

  5. nelson-stack nelson-stack Public

    AI safety audit, local-first tooling, and evidence-backed engineering — the connective tissue across Nelson's repos.

    Python

  6. nullbench nullbench Public

    Pre-register decisions. Score them against chance. Never backfill.

    Python