A GGML/C++ engine for the neural-network front half of TencentARC/FreeSplatter (image patch tokenizer → multi-view self-attention transformer → per-pixel 3D-Gaussian parameter head) plus the downstream camera-pose recovery and cross-run registration that consume its output. Given N uncalibrated views it returns, per input pixel, the activated Gaussian parameters; PnP then recovers each view's camera, and successive runs are aligned into one accumulating world — the path toward live reconstruction from a moving camera.
Scope (updated): the engine (pieces 1–3), PnP pose recovery (now IN
scope), and the cross-run Sim(3) alignment / accumulation. Keep the seam at the
[N, H, W, gaussian_channels] tensor clean — it is the contract between the
engine and the pose consumer. Rendering itself stays in the demo viewer
(Vulkan/WebGL), not the engine.
Target checkpoint first: freesplatter-scene (gaussian_channels=23,
sh_residual=true, black background). The transformer backbone is identical
across all three variants, so object/object-2dgs are cheap follow-ons.
- Everything ships in C++ (the engine, PnP, and the alignment/accumulation once proven). Go only for the demo web server — the purego layer that drives the C API → Vulkan + WebGL viewer.
- The CLI and the C API must have NO Python dependency at runtime. Every piece
of current and future functionality is reachable from
free_splatter-cliandinclude/free_splatter.hwithout invoking Python. - Python is confined to two places, neither shipped:
- Dev-time reference / conversion / validation that runs in
docker/Dockerfile.cuda(scripts/hf_dump.py,convert.py,compare_taps.py, …) — the only place torch runs; never a runtime dependency. - The
pose/research prototype — now DONE and DELETED. It was the temporary Python (numpy + cv2) prototype that proved the accumulating-reconstruction approach. That approach is proven, the whole pipeline (focal, Sim(3) align, robust PnP, accumulation, loop closure, consensus fusion) is rewritten in C++ (src/pose.{h,cpp}), exposed viafree_splatter-cli+include/free_splatter.h, and the Python prototype has been removed (see git history for it and its layer-by-layer parity harnesses). No Python remains in the pose path.
- Dev-time reference / conversion / validation that runs in
This is a numerical port. Correctness means matching the PyTorch reference layer by layer, not "it runs and looks plausible."
- Every op is tapped. Each meaningful intermediate gets a stable
ggml_set_name()and is dumped (free_splatter-cli --dump-taps DIR) in the formatscripts/compare_taps.pyreads. - A piece is not "done" until its taps pass
compare_taps.pyagainst the float64 PyTorch reference (scripts/hf_dump.py), in topological order, with the first divergence fixed — never paper over a downstream symptom of an upstream bug. compare_taps stops at the first failing tap on purpose. - New graph ops land with an asset-free golden-op test first
(
tests/test_graph_blocks.cpp), pinned to hand-computed f64 references, before any fixture-based parity claim. A wrong op should fail with zero fixtures. - Gates: per-row
cosine ≥ 0.99999AND normalizedrow-err ≤ 1e-2on CPU-f32. Widen only viacompare_taps.py --scale Nfor Vulkan/fp16, and write down why N. - Sensitive ops are computed in f64 in the reference (attention softmax, GELU, norms): the engine targets the exact value, not a second noisy float path. Bring each stage up in f32 first to isolate algorithmic bugs, then switch weights to f16 and re-check the head error stays within gate.
- Known silent-bug hotspots (test these in isolation before trusting
anything downstream): the patchify
(ic,kh,kw)flatten order; the unpatchify(p,q,c)pixel-unshuffle order; attention precision (useggml_flash_attn_ext/f32, never f16 softmax); GELU must beggml_gelu_erf(exact erf), neverggml_gelu(tanh); the scale activation (min + (max-min)·sigmoid, not exp).
Each component (image, gguf_loader, backend, model, the head, and
pose = PnP + focal + Sim(3) alignment + accumulation/loop/fusion) has its own
unit test and is made independently green before cross-component parity.
Component boundaries match the file layout. Keep the seam at the
[N,H,W,gaussian_channels] tensor clean — that is the contract between the engine
and the pose/rendering consumers. The C++ pose component (tests/test_pose.cpp,
asset-free golden tier) carries the parity discipline the Python prototype
established: bit-exact to upstream estimate_poses and validated against
independent ground-truth poses — see git history for the prototype's
check_upstream_parity.py / re10k_experiment.py harnesses.
- Collect information first. On any bug, add/inspect taps, profile, trace shapes and dtypes, and survey the upstream FreeSplatter source / the relevant papers / other ggml ports before proposing a fix. Make the cause unambiguous first.
- Prefer admitting ignorance and gathering more context over committing to a fix on a hunch. A wrong fix that masks a symptom is worse than no fix.
- Verify the fix did what you predicted. After fixing, re-run the specific failing tap and confirm it now passes for the reason you expected — check that neighbouring taps did not shift, and that the gate which was failing is the gate now passing. Do not declare victory on a green aggregate.
- The GGUF model file is TRUSTED — not fuzzed, not hardened. Loading is for our own converted weights.
- Image inputs are UNTRUSTED — validated in
src/image.cpp, fuzzed (fuzz/fuzz_image.cpp), and the test suite runs ASan/UBSan-clean. Keep this asymmetry intentional.
nix developthencmake --preset debug && cmake --build --preset debug— ASan/UBSan, the verification build.ctest --preset debug -LE model— fast, asset-free tier (golden-op + unit).FREE_SPLATTER_GGUF=…f32.gguf FREE_SPLATTER_FIXTURES=…/2view ctest -L model— full per-layer + head parity (needs converted weights + fixtures;FREE_SPLATTER_MAX_BLOCKS=1for a fast piece-1 + block-0 subset).cmake --preset vulkan— Vulkan backend parity.docker/Dockerfile.cudais the only place the upstream PyTorch model runs (reference taps + weight conversion). The engine never depends on torch/CUDA.
The README is for the end user with the least interest in internals. Keep technical and validation detail HERE (and in script docstrings), not in the README. See its top comment for the audience rule.