Skip to content

docs: publish GLM-5.2 inference research closeout - #794

Open
ZacharyZcR wants to merge 1 commit into
JustVugg:devfrom
ZacharyZcR:docs/glm52-inference-research-closeout
Open

docs: publish GLM-5.2 inference research closeout#794
ZacharyZcR wants to merge 1 commit into
JustVugg:devfrom
ZacharyZcR:docs/glm52-inference-research-closeout

Conversation

@ZacharyZcR

Copy link
Copy Markdown
Contributor

Summary

  • publish the controlled inference-paper experiment matrix and failure ledger
  • add reproducible placement, batching, prefix reuse, interference, disaggregation, and entropy tools
  • retain fixed replay fixtures and the measured lossless split analysis

Key conclusion

For lossless single-stream GLM-5.2 decode on the tested 6x RTX 5090 / dual Xeon host, the remaining engine-only headroom is incremental. Future claims should reduce measured bytes per token or remove a measured whole-layer critical-path cost under held-out end-to-end evaluation.

Validation

  • make -C c check: 292 Python tests passed, 18 skipped; native suite passed
  • python3 -m unittest discover -s tests -p "test_placement_*.py": 8 passed
  • CLI smoke checks for the included Python tools

Related: #537, #594, #452

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant