diff --git a/README.md b/README.md
index 55107539f..f0feefc83 100644
--- a/README.md
+++ b/README.md
@@ -64,6 +64,10 @@ to load, even when their model family name looks compatible.
The current production contract is Apple Silicon + Metal + embedded
ExpertMajor v2. AUTO is the normal startup mode.
+> Published artifact filenames carrying the legacy prefix are frozen compatibility
+> identifiers tied to existing hashes and manifests. They do not represent the
+> current product name; the runtime and user-facing project are Hebrus.
+
| Model family | Minimum unified memory | Qualified execution | Artifact availability |
| --- | ---: | --- | --- |
| DeepSeek V4 Flash | 64 GiB | AUTO resolving to resident or SSD; explicit modes for qualification | Published ExpertMajor v2 artifact; `download_model.sh deepseek-v2` |
@@ -263,12 +267,11 @@ can create substantial memory pressure and I/O. Follow AUTO, use only exact
qualified artifacts, monitor swap and memory pressure, and do not expose a
local server to untrusted networks without an appropriate security boundary.
-The separate companion desktop application is **Hebrus Studio**, currently
-developed in the existing
-[`andreaborio/dsbox`](https://github.com/andreaborio/dsbox) fork until its
-administrative rename. Its bridge release preserves the DSBox bundle ID, data
-locations, environment namespace, storage keys, and rollback path so the public
-rename does not split existing installations.
+The companion desktop application is published as
+[`andreaborio/hebrus-studio`](https://github.com/andreaborio/hebrus-studio).
+Its compatibility bridge preserves the legacy DSBox bundle ID, data locations,
+environment namespace, storage keys, and rollback path so the public rename
+does not split existing installations.
## Contributing and security
diff --git a/dir-steering/README.md b/dir-steering/README.md
index e1fdbfe5a..409d5a130 100644
--- a/dir-steering/README.md
+++ b/dir-steering/README.md
@@ -1,15 +1,15 @@
# Directional Steering
-Directional steering is a runtime activation edit for DS4. A steering file is a
+Directional steering is a runtime activation edit for Hebrus. A steering file is a
flat `f32` matrix with one normalized 4096-wide direction per layer. During
-inference, ds4 can apply the edit after attention outputs, FFN outputs, or both:
+inference, Hebrus can apply the edit after attention outputs, FFN outputs, or both:
```text
y = y - scale * direction[layer] * dot(direction[layer], y)
```
Positive scale removes the represented direction. Negative scale amplifies it.
-With no steering file or zero scales, ds4 follows the normal inference path.
+With no steering file or zero scales, Hebrus follows the normal inference path.
## Runtime Options
@@ -39,8 +39,8 @@ Build the vector:
```sh
python3 dir-steering/tools/build_direction.py \
- --ds4 ./ds4 \
- --model ds4flash.gguf \
+ --ds4 ./hebrus \
+ --model deepseek-v4-flash.gguf \
--good-file dir-steering/examples/succinct.txt \
--bad-file dir-steering/examples/verbose.txt \
--out dir-steering/out/verbosity.json \
@@ -58,7 +58,7 @@ dir-steering/out/verbosity.f32
Try a terse run:
```sh
-./ds4 -m ds4flash.gguf --nothink --temp 0 -n 160 \
+./hebrus -m deepseek-v4-flash.gguf --nothink --temp 0 -n 160 \
--dir-steering-file dir-steering/out/verbosity.f32 \
--dir-steering-ffn -1 \
-p "Explain why databases use indexes."
@@ -67,7 +67,7 @@ Try a terse run:
Try a verbose run:
```sh
-./ds4 -m ds4flash.gguf --nothink --temp 0 -n 220 \
+./hebrus -m deepseek-v4-flash.gguf --nothink --temp 0 -n 220 \
--dir-steering-file dir-steering/out/verbosity.f32 \
--dir-steering-ffn 2 \
-p "Explain why databases use indexes."
@@ -85,8 +85,8 @@ Use the sweep helper to test several strengths on a fixed prompt set:
```sh
python3 dir-steering/tools/run_sweep.py \
- --ds4 ./ds4 \
- --model ds4flash.gguf \
+ --ds4 ./hebrus \
+ --model deepseek-v4-flash.gguf \
--direction dir-steering/out/verbosity.f32 \
--prompts dir-steering/examples/eval_prompts.txt \
--scales "-1,-0.5,0,0.5,1,2" \
@@ -124,7 +124,7 @@ The extractor compares two prompt sets:
- `good-file`: target prompts for the direction you want to represent.
- `bad-file`: contrast prompts that should be separated from the target.
-It captures DS4 activations from the same local GPU graph used for inference,
+It captures Hebrus activations from the same local GPU graph used for inference,
averages target minus contrast, normalizes one vector per layer, and writes both
metadata JSON and the runtime `.f32` file.
diff --git a/docs/architecture/hebrus-runtime-flow.svg b/docs/architecture/hebrus-runtime-flow.svg
index 79ca25d0e..b43ca648d 100644
--- a/docs/architecture/hebrus-runtime-flow.svg
+++ b/docs/architecture/hebrus-runtime-flow.svg
@@ -78,7 +78,7 @@
• index routed-expert
records before inference
- ds4.expert_major.v2
+ ExpertMajor v2
diff --git a/gguf-tools/README.md b/gguf-tools/README.md
index f9ef08832..72e6806e2 100644
--- a/gguf-tools/README.md
+++ b/gguf-tools/README.md
@@ -1,19 +1,19 @@
-# DS4 GGUF Tools
+# Hebrus GGUF Tools
This directory contains the offline tools used to build and evaluate the
-DeepSeek V4, GLM 5.2, and Qwen3.6 GGUF files supported by `ds4`.
+DeepSeek V4, GLM 5.2, and Qwen3.6 GGUF files supported by Hebrus.
The important pieces are:
- `deepseek4-quantize.c`: C HF-safetensors to GGUF quantizer.
- `quants.[ch]`: the deliberately small local quantization implementation used
- by the quantizer. It implements the DS4 output formats we actually ship:
+ by the quantizer. It implements the Hebrus output formats we actually ship:
`q8_0`, `q4_K`, `q2_K`, and `iq2_xxs`.
- `ds4-expert-major.py`: deterministic canonical-to-native layout converter
and byte-level verifier for DeepSeek, GLM, and Qwen
`ds4.expert_major.v2` GGUFs.
- `imatrix/`: dataset and instructions for collecting routed-MoE activation
- importance with `ds4`.
+ importance with Hebrus.
- `quality-testing/`: prompts and scripts used to compare local GGUF variants
against official DeepSeek V4 Flash continuations.
@@ -27,7 +27,7 @@ The quantizer is plain C and does not link GGML. GGUF metadata handling,
safetensors loading, FP4/FP8 dequantization, and the quantizers used by our Q2
and Q4 recipes live in this directory.
-## Build a DS4-native ExpertMajor v2 GGUF
+## Build a Hebrus ExpertMajor v2 GGUF
Reorder an already qualified DeepSeek, GLM, or Qwen GGUF without changing
quantization:
@@ -60,7 +60,7 @@ filesystem.
Run a completed artifact without conversion flags:
```sh
-./ds4 -m /absolute/path/to/MODEL-DS4-ExpertMajor-v2.gguf --ctx 8192
+./hebrus -m /absolute/path/to/MODEL-DS4-ExpertMajor-v2.gguf --ctx 8192
```
## Generate An Imatrix
@@ -71,22 +71,22 @@ First regenerate or inspect the calibration dataset:
python3 gguf-tools/imatrix/dataset/build_ds4_imatrix_dataset.py
```
-Then collect activation statistics with the DS4 runtime:
+Then collect activation statistics with the Hebrus runtime:
```sh
-./ds4 \
+./hebrus \
-m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
--imatrix-dataset gguf-tools/imatrix/dataset/rendered_prompts.txt \
--imatrix-out gguf/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4.dat \
--ctx 32768
```
-The imatrix file is useful immediately with this DS4 quantizer. Generic GGUF
-tools need DS4-specific tensor-name mapping and per-expert slicing before they
+The imatrix file is useful immediately with the Hebrus quantizer. Generic GGUF
+tools need Hebrus-specific tensor-name mapping and per-expert slicing before they
can use it correctly. The accepted imatrix format is the legacy llama.cpp
-binary `.dat` file emitted by `ds4 --imatrix-out`.
+binary `.dat` file emitted by `hebrus --imatrix-out`.
-Generating this `.dat` file locally is possible, but slow: it runs the DS4
+Generating this `.dat` file locally is possible, but slow: it runs the Hebrus
prefill graph over the full calibration corpus and reads routed-MoE activation
statistics back from the GPU. The latest published imatrix-generated GGUF files
are available in the antirez Hugging Face repository:
@@ -157,7 +157,7 @@ from the dequantized weight itself:
importance[column] = sum(row[column]^2) over all rows
```
-This is a weight-energy heuristic. It is not as good as measuring real DS4
+This is a weight-energy heuristic. It is not as good as measuring real Hebrus
activations, but it gives the quantizer a stable column weighting and was good
enough for the first working 2-bit GGUFs.
diff --git a/gguf-tools/imatrix/README.md b/gguf-tools/imatrix/README.md
index 15109160e..65b87df05 100644
--- a/gguf-tools/imatrix/README.md
+++ b/gguf-tools/imatrix/README.md
@@ -1,4 +1,4 @@
-# DS4 Imatrix Pipeline
+# Hebrus Imatrix Pipeline
This directory contains the calibration dataset and instructions used to build
activation importance matrices for DeepSeek V4 Flash and Pro GGUF
@@ -15,7 +15,7 @@ layer. Both variants expose three routed expert tensors per layer:
For gate/up tensors, the collector records the squared FFN-normalized input
activation. For down tensors, it records the squared routed SwiGLU row after
route weighting. The result tells the quantizer which input columns are used
-more heavily by the actual DS4 inference graph.
+more heavily by the actual Hebrus inference graph.
## 1. Build The Calibration Dataset
@@ -33,14 +33,14 @@ The important output is:
gguf-tools/imatrix/dataset/rendered_prompts.txt
```
-It contains DS4-rendered chat prompts, separated by visible
+It contains Hebrus-rendered chat prompts, separated by visible
`DS4_IMATRIX_PROMPT` markers. The prompts include:
- C/Metal source-review prompts from this repository.
- Long-context snippets.
-- Agent/tool-call prompts using DS4's DSML syntax.
+- Agent/tool-call prompts using Hebrus' DSML syntax.
- Language/prose rewriting, summarization, extraction, and translation prompts.
-- `ds4-eval` GPQA Diamond, SuperGPQA, and AIME2025 benchmark prompts.
+- `hebrus-eval` GPQA Diamond, SuperGPQA, and AIME2025 benchmark prompts.
- Both thinking and non-thinking assistant prefixes.
Check the generated `gguf-tools/imatrix/dataset/manifest.json` for the exact
@@ -48,7 +48,7 @@ record count and coarse token estimate of the current tree.
## 2. Collect The Imatrix
-Use the DS4 runtime itself to collect routed MoE activation statistics. The
+Use the Hebrus runtime itself to collect routed MoE activation statistics. The
collector uses the loaded GGUF metadata. Runtime admission still applies: the
collector must use an admitted ExpertMajor v2 artifact, not a canonical
converter input.
@@ -56,7 +56,7 @@ converter input.
Flash example:
```sh
-./ds4 \
+./hebrus \
-m DEEPSEEK-DS4-ExpertMajor-v2.gguf \
--imatrix-dataset gguf-tools/imatrix/dataset/rendered_prompts.txt \
--imatrix-out ../deepseek-v4-quants/imatrix/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4-1p5m.dat \
@@ -70,10 +70,10 @@ an executable collection recipe.
Useful smoke-test limits:
```sh
-./ds4 \
+./hebrus \
-m DEEPSEEK-DS4-ExpertMajor-v2.gguf \
--imatrix-dataset gguf-tools/imatrix/dataset/rendered_prompts.txt \
- --imatrix-out /tmp/ds4-test.imatrix.dat \
+ --imatrix-out /tmp/hebrus-test.imatrix.dat \
--imatrix-max-prompts 1 \
--imatrix-max-tokens 4096
```
@@ -82,7 +82,7 @@ The collector is Metal-only because it hooks the layer-major Metal prefill graph
It does not change inference math; it reads the already materialized MoE inputs
and accumulates `sum(x[column]^2)` per routed expert.
-The output format is llama.cpp's legacy binary `.dat` imatrix format. DS4 packs
+The output format is llama.cpp's legacy binary `.dat` imatrix format. Hebrus packs
per-expert vectors into one entry per routed expert tensor:
```text
@@ -170,11 +170,11 @@ avg greedy LCP: 12.21 imatrix / 11.94 old
## Compatibility
The `.dat` file is intentionally in llama.cpp's legacy imatrix format, so the
-data is not conceptually tied to DS4. In practice, it is immediately useful
-only with a quantizer that understands DS4's tensor names and packed per-expert
+data is not conceptually tied to Hebrus. In practice, it is immediately useful
+only with a quantizer that understands Hebrus' tensor names and packed per-expert
entries. The current `deepseek4-quantize` tooling does that.
Other GGUF creation tools can use the same imatrix if they implement the same
-name mapping and per-expert slicing convention. Without that DS4-specific
+name mapping and per-expert slicing convention. Without that Hebrus-specific
mapping, a generic imatrix loader will see valid data but will not know how to
apply the packed routed-expert vectors correctly.
diff --git a/gguf-tools/imatrix/dataset/README.md b/gguf-tools/imatrix/dataset/README.md
index c686e92eb..e47a5ce25 100644
--- a/gguf-tools/imatrix/dataset/README.md
+++ b/gguf-tools/imatrix/dataset/README.md
@@ -1,6 +1,6 @@
-# DS4 Imatrix Calibration Dataset
+# Hebrus Imatrix Calibration Dataset
-This directory contains DS4-rendered chat prompts for collecting activation
+This directory contains Hebrus-rendered chat prompts for collecting activation
statistics before building new low-bit GGUF files.
Run:
@@ -36,7 +36,7 @@ Some records include DSML tool schemas, sampled DSML tool calls, and tool-result
turns so the imatrix sees the same special-token patterns used by agent clients.
The corpus is provider-neutral and also includes language/prose rewriting,
summarization, copy-editing, extraction, multilingual translation, programming
-prompts, Bash scripting, algorithm recall, `ds4-eval` benchmark-reasoning
+prompts, Bash scripting, algorithm recall, `hebrus-eval` benchmark-reasoning
prompts, long-context code synthesis, agent transcript replay, log diagnosis,
prose fact recovery, delayed-constraint and small needle tasks, Metal/C code
review tasks, and inference-specific debugging tasks.
diff --git a/gguf-tools/quality-testing/README.md b/gguf-tools/quality-testing/README.md
index 10b2ee6e2..6148ae820 100644
--- a/gguf-tools/quality-testing/README.md
+++ b/gguf-tools/quality-testing/README.md
@@ -48,7 +48,7 @@ tracked because they are derived from an external API.
make -C gguf-tools quality-score
```
-The scorer links against the DS4 runtime and uses Metal by default.
+The scorer links against the Hebrus runtime and uses Metal by default.
## 3. Score GGUF Variants
diff --git a/speed-bench/README.md b/speed-bench/README.md
index 61f182211..35e76746c 100644
--- a/speed-bench/README.md
+++ b/speed-bench/README.md
@@ -8,13 +8,13 @@ Run each retained frontier in a separate process. For example:
```
for frontier in 128 2048 8192 32768; do
- ./ds4-bench \
- -m /absolute/path/to/QUALIFIED-DS4-ExpertMajor-v2.gguf \
+ ./hebrus-bench \
+ -m /absolute/path/to/qualified-model.gguf \
--prompt-file speed-bench/promessi_sposi.txt \
--ctx-start "$frontier" --ctx-max "$frontier" \
--ctx-alloc 65536 --gen-tokens 128 \
- --csv "/tmp/ds4-$frontier.csv" \
- --dump-decode-evidence-dir "/tmp/ds4-$frontier-evidence"
+ --csv "/tmp/hebrus-$frontier.csv" \
+ --dump-decode-evidence-dir "/tmp/hebrus-$frontier-evidence"
done
```
@@ -27,7 +27,7 @@ and 100,000-token lanes defined in `CONTRIBUTING.md`.
On the qualified M5 lane, `run_m5_dsflash_arm.sh` extends an undersized prompt
deterministically for frontiers above 32,768 tokens. The generated prompt, its
source, both SHA-256 values, and whether extension occurred are recorded beside
-the run. `ds4-bench` still verifies the actual model-specific token count and
+the run. `hebrus-bench` still verifies the actual model-specific token count and
fails closed if the prompt is too short.
Provide PR including your numbers if your hardware was not already tested.
diff --git a/tests/qwen/README.md b/tests/qwen/README.md
index f379159ea..730065a1f 100644
--- a/tests/qwen/README.md
+++ b/tests/qwen/README.md
@@ -1,5 +1,8 @@
# Qwen3.6 fixtures and ExpertMajor v2 gates
+The internal `ds4_test` target retains its historical compatibility name;
+runtime examples use the public `hebrus` executable.
+
This directory contains frozen Qwen3.6 reference data, reproducible fixture
collectors, and narrow offline checks. Current production inference accepts
only the qualified embedded ExpertMajor v2 artifact on Apple Metal. The support
@@ -141,7 +144,7 @@ table above. Normal startup is flag-free AUTO:
```sh
QWEN_V2=/absolute/path/to/Qwen3.6-35B-A3B-DS4-ExpertMajor-v2-MLX-Affine4-G64.gguf
-./ds4 -m "$QWEN_V2" --ctx 8192 \
+./hebrus -m "$QWEN_V2" --ctx 8192 \
-n 32 --temp 0 \
-p 'Scrivi solo una breve funzione Python che somma due interi.'
```
@@ -166,17 +169,17 @@ a host where resident admission succeeds, run the same deterministic prompt in
both modes:
```sh
-./ds4 -m "$QWEN_V2" --ctx 8192 --resident \
+./hebrus -m "$QWEN_V2" --ctx 8192 --resident \
-n 32 --temp 0 \
-p 'Scrivi solo una breve funzione Python che somma due interi.'
-./ds4 -m "$QWEN_V2" --ctx 8192 --ssd-streaming \
+./hebrus -m "$QWEN_V2" --ctx 8192 --ssd-streaming \
-n 32 --temp 0 \
-p 'Scrivi solo una breve funzione Python che somma due interi.'
```
Compare deterministic output and logits, not plausibility alone. Resident mode
-must use complete mapped tensors with zero DS4 expert-cache `pread` accounting.
+must use complete mapped tensors with zero Hebrus expert-cache `pread` accounting.
SSD mode must allocate the first 321-expert slab within its admitted budget and
must not introduce swap. Keep warm page-cache evidence separate from cold
device-I/O evidence, and never bypass a failed admission to obtain a number.
diff --git a/tests/test-vectors/README.md b/tests/test-vectors/README.md
index de3acaace..f20415bd5 100644
--- a/tests/test-vectors/README.md
+++ b/tests/test-vectors/README.md
@@ -1,5 +1,9 @@
# DeepSeek V4 Flash Test Vectors
+The internal `ds4_test` binary and `DS4_*` environment variables retain their
+historical names as test and compatibility interfaces. Public runtime examples
+use the `hebrus` executable.
+
These vectors were captured from the official DeepSeek V4 Flash API using
`deepseek-v4-flash`, greedy decoding, thinking disabled, and
`top_logprobs=20`. The hosted API does not expose full logits, so these files
@@ -10,7 +14,7 @@ Files:
- `prompts/*.txt`: exact user prompts.
- `official/*.official.json`: official API continuations and top-logprobs.
- `official.vec`: compact C-test fixture generated from the official JSON.
-- `local-golden.vec`: local top-k/logit fixture captured from a known-sane DS4
+- `local-golden.vec`: local top-k/logit fixture captured from a known-sane Hebrus
Flash run. It is used to catch substantial backend drift that can keep the
same greedy token while damaging the logits distribution.
@@ -57,8 +61,8 @@ API response.
To inspect a local top-logprob dump manually:
```sh
-./ds4 --metal --nothink -sys "" --temp 0 -n 4 --ctx 16384 \
+./hebrus --metal --nothink -sys "" --temp 0 -n 4 --ctx 16384 \
--prompt-file tests/test-vectors/prompts/long_code_audit.txt \
- --dump-logprobs /tmp/long_code_audit.ds4.json \
+ --dump-logprobs /tmp/long_code_audit.hebrus.json \
--logprobs-top-k 20
```