Jacobian-Brainwash : A manual alignment tool for large language models built on Anthropic's Jacobian Lens. Results are exportable.
-
Updated
Aug 1, 2026 - Python
Jacobian-Brainwash : A manual alignment tool for large language models built on Anthropic's Jacobian Lens. Results are exportable.
Preregistered Jacobian-lens reliability research on Gemma 4: 25k prompts, frozen probes, public traces, a prospective transfer miss, and an answer-identity confound.
OpenAI-compatible server with a live view into any HF model's residual stream. pip install brainscope
Agents that talk through model internals — activations & J-space — instead of text. No words pass between them. The lab on top of brainscope + hidden-directions.
LLM decision boundary; LLM reasoning trajectory; Geometric Lens; Laguerre Geometry
EVict and recOver KV cache Entries. Selective KV cache eviction and recovery for long-context LLM inference.
A Jacobian-Lens (J-Lens) observer for vision-language models — read what a VLM is poised to say, before it says it. Multimodal J-Lens on Qwen3.5, concept-race, and a forward-only prompt helper.
Anthropic's Jacobian Lens, ported to Mac. The official reference implementation targets CUDA GPUs — this fork runs it on Apple silicon (PyTorch MPS) and adds a local interactive explorer for Qwen with plain-language interpretability demos and a live cause-and-effect intervention.
Watch a language model's thoughts form before it speaks — interactive workbench + findings for Anthropic's Jacobian lens
Jacobian Lens (J-Lens) and J-Space Toolkit for transformer mechanistic interpretability. Train linear lenses, decompose hidden states, and run causal interventions on decoder-only models.
Probing LLM J-spaces through the Jacobian lens on one RTX 3090 — 421-record research data dump (Units 0–15, incl. workspace-span batteries) with live dashboard, films, and per-record commentary
Anthropic's jlens for single or multi-model chat replay, with a web interface.
J-lens for audio-input LLMs with reproducible mixed text and audio fitting
Independent interpretability research (DRAFT — not yet run): is an inferred character trait stored, or reconstructed from the scene on demand? Across model scale + a scene-masked direction probe
Autonomous J-space search: causal concept-swap interventions inside a language model, built on the Jacobian lens (Anthropic, 2026).
Local-first macOS desktop application for fitting, applying, and exploring Jacobian Lenses on open-weight decoder language models.
Pre-registered Jacobian-lens probe: does a verbatim-stated trait ("Maria is generous") persist in a model's J-space differently than an inferred one, as neutral text intervenes? Gema-3-4B-IT.
Jacobian Lens experiments for language model interpretability.
Workspace-lens audit cards and training-workflow scaffolds for open-weight language models.
Chain-of-density study of 'Verbalizable Representations Form a Global Workspace in Language Models' (Gurnee, Sofroniew, Lindsey et al. 2026, Transformer Circuits/Anthropic) - five-tier note, locator-verified. The Jacobian lens, the J-space, alignment auditing, counterfactual reflection training.
Add a description, image, and links to the jacobian-lens topic page so that developers can more easily learn about it.
To associate your repository with the jacobian-lens topic, visit your repo's landing page and select "manage topics."