diff --git a/README.md b/README.md
index d1e96f1..a1e530e 100644
--- a/README.md
+++ b/README.md
@@ -24,7 +24,7 @@ A collection of enhancements, plugins, and prompts for [open-webui](https://gith
| Rank | Plugin | Version | Downloads | Views | 📅 Updated |
| :---: | :--- | :---: | :---: | :---: | :---: |
| 🥇 | [Smart Mind Map](https://openwebui.com/posts/turn_any_text_into_beautiful_mind_maps_3094c59a) |  |  |  |  |
-| 🥈 | [Async Context Compression](https://openwebui.com/posts/async_context_compression_b1655bc8) |  |  |  |  |
+| 🥈 | [Async Context Compression](https://openwebui.com/posts/async_context_compression_b1655bc8) |  |  |  |  |
| 🥉 | [Smart Infographic](https://openwebui.com/posts/smart_infographic_ad6f0c7f) |  |  |  |  |
| 4️⃣ | [Markdown Normalizer](https://openwebui.com/posts/markdown_normalizer_baaa8732) |  |  |  |  |
| 5️⃣ | [OpenWebUI Skills Manager Tool](https://openwebui.com/posts/openwebui_skills_manager_tool_b4bce8e4) |  |  |  |  |
diff --git a/README_CN.md b/README_CN.md
index e89cc41..6357634 100644
--- a/README_CN.md
+++ b/README_CN.md
@@ -21,7 +21,7 @@ OpenWebUI 增强功能集合。包含个人开发与收集的插件、提示词
| 排名 | 插件 | 版本 | 下载 | 浏览 | 📅 更新 |
| :---: | :--- | :---: | :---: | :---: | :---: |
| 🥇 | [Smart Mind Map](https://openwebui.com/posts/turn_any_text_into_beautiful_mind_maps_3094c59a) |  |  |  |  |
-| 🥈 | [Async Context Compression](https://openwebui.com/posts/async_context_compression_b1655bc8) |  |  |  |  |
+| 🥈 | [Async Context Compression](https://openwebui.com/posts/async_context_compression_b1655bc8) |  |  |  |  |
| 🥉 | [Smart Infographic](https://openwebui.com/posts/smart_infographic_ad6f0c7f) |  |  |  |  |
| 4️⃣ | [Markdown Normalizer](https://openwebui.com/posts/markdown_normalizer_baaa8732) |  |  |  |  |
| 5️⃣ | [OpenWebUI Skills Manager Tool](https://openwebui.com/posts/openwebui_skills_manager_tool_b4bce8e4) |  |  |  |  |
diff --git a/docs/community-stats.json b/docs/community-stats.json
index 97fb520..2f9d0cd 100644
--- a/docs/community-stats.json
+++ b/docs/community-stats.json
@@ -36,7 +36,7 @@
"title": "Async Context Compression",
"slug": "async_context_compression_b1655bc8",
"type": "filter",
- "version": "1.7.0",
+ "version": "1.7.1",
"author": "Fu-Jie",
"description": "Reduces token consumption in long conversations while maintaining coherence through intelligent summarization and message compression.",
"downloads": 2893,
diff --git a/docs/community-stats.md b/docs/community-stats.md
index 2c0d5df..29551fd 100644
--- a/docs/community-stats.md
+++ b/docs/community-stats.md
@@ -37,7 +37,7 @@
| Rank | Title | Type | Version | Downloads | Views | Upvotes | Saves | Updated |
|:---:|------|:---:|:---:|:---:|:---:|:---:|:---:|:---:|
| 1 | [Smart Mind Map](https://openwebui.com/posts/turn_any_text_into_beautiful_mind_maps_3094c59a) | action |  |  |  |  |  | 2026-04-24 |
-| 2 | [Async Context Compression](https://openwebui.com/posts/async_context_compression_b1655bc8) | filter |  |  |  |  |  | 2026-05-22 |
+| 2 | [Async Context Compression](https://openwebui.com/posts/async_context_compression_b1655bc8) | filter |  |  |  |  |  | 2026-05-22 |
| 3 | [Smart Infographic](https://openwebui.com/posts/smart_infographic_ad6f0c7f) | action |  |  |  |  |  | 2026-04-25 |
| 4 | [Markdown Normalizer](https://openwebui.com/posts/markdown_normalizer_baaa8732) | filter |  |  |  |  |  | 2026-04-15 |
| 5 | [OpenWebUI Skills Manager Tool](https://openwebui.com/posts/openwebui_skills_manager_tool_b4bce8e4) | tool |  |  |  |  |  | 2026-05-22 |
diff --git a/docs/community-stats.zh.md b/docs/community-stats.zh.md
index 7f7ad2c..6447be7 100644
--- a/docs/community-stats.zh.md
+++ b/docs/community-stats.zh.md
@@ -37,7 +37,7 @@
| 排名 | 标题 | 类型 | 版本 | 下载 | 浏览 | 点赞 | 收藏 | 更新日期 |
|:---:|------|:---:|:---:|:---:|:---:|:---:|:---:|:---:|
| 1 | [Smart Mind Map](https://openwebui.com/posts/turn_any_text_into_beautiful_mind_maps_3094c59a) | action |  |  |  |  |  | 2026-04-24 |
-| 2 | [Async Context Compression](https://openwebui.com/posts/async_context_compression_b1655bc8) | filter |  |  |  |  |  | 2026-05-22 |
+| 2 | [Async Context Compression](https://openwebui.com/posts/async_context_compression_b1655bc8) | filter |  |  |  |  |  | 2026-05-22 |
| 3 | [Smart Infographic](https://openwebui.com/posts/smart_infographic_ad6f0c7f) | action |  |  |  |  |  | 2026-04-25 |
| 4 | [Markdown Normalizer](https://openwebui.com/posts/markdown_normalizer_baaa8732) | filter |  |  |  |  |  | 2026-04-15 |
| 5 | [OpenWebUI Skills Manager Tool](https://openwebui.com/posts/openwebui_skills_manager_tool_b4bce8e4) | tool |  |  |  |  |  | 2026-05-22 |
diff --git a/docs/development/async-context-compression-idless-summary-reuse-review-gpt-5.5.zh.md b/docs/development/async-context-compression-idless-summary-reuse-review-gpt-5.5.zh.md
new file mode 100644
index 0000000..6d51b2e
--- /dev/null
+++ b/docs/development/async-context-compression-idless-summary-reuse-review-gpt-5.5.zh.md
@@ -0,0 +1,78 @@
+---
+title: "代码评审:Async Context Compression idless request summary reuse"
+type: review
+status: addressed
+date: 2026-06-26
+target_commit: 07d26c1a1e641f6b79f20d199e8b807c27d048ee
+target: plugins/filters/async-context-compression
+reviewer: gpt-5.5
+---
+
+# 代码评审:Async Context Compression idless request summary reuse
+
+## Code Review Results
+
+**Scope:** `extensions` 子仓库最新提交 `07d26c1 fix(async-context-compression): reuse summaries for idless requests`。
+
+**Intent:** 当 Open WebUI middleware 发给模型的 request body 缺少稳定 message ids,或 DB active branch 比 request body 多一个 terminal assistant placeholder 时,插件先用 DB active branch 证明 body 与当前可见分支一致,再复用已有 branch-aware summary,避免错误退回超长原始历史。
+
+**Mode:** report-only review。主线程复核代码并运行插件测试;并行使用 correctness、testing、maintainability 评审视角。未修改实现代码。
+
+**Reviewers:** correctness、testing、maintainability、主线程复核。
+
+## Addressing Update
+
+2026-06-26 已处理本轮 review 的所有 actionable 项:
+
+- P2 #1: debug 日志路径改为使用当前 snapshot 实际使用的 `current_refs_for_snapshot`,不再在 idless fallback 场景对 `current_refs=None` 切片。新增 `test_snapshot_selection_debug_handles_idless_tail_with_multiple_snapshots` 覆盖 `debug_mode=True + 多 snapshot + idless tail`。
+- P2 #2: `_unfold_db_branch_for_body_ref_fallback()` 现在捕获 `convert_output_to_messages()` 的普通异常并 fail closed,拒绝 DB ref fallback,而不是让 inlet 失败。新增 `test_unfold_db_branch_fallback_rejects_conversion_errors`。
+- P2 #3: 新增 `test_inlet_reuses_same_length_idless_body_that_omits_db_output`,覆盖同长度 idless body 省略 DB assistant `output` 的正向复用,并断言 tool call mismatch 会拒绝 fallback。
+- P2 #4: 新增 `test_load_full_chat_messages_filters_failed_assistant_from_direct_messages`,覆盖 direct `chat["messages"]` 路径的 failed assistant 过滤。
+- P3 #5: 新增 `test_unfolded_db_message_allows_trimmed_assistant_only_with_metadata`,覆盖 assistant `metadata.tool_outputs_trimmed` 容忍分支及缺少 metadata 时的负例。
+
+验证:
+
+- `../.venv/bin/python -m pytest plugins/filters/async-context-compression/test_async_context_compression.py -k "idless or output or failed_assistant or tool_outputs_trimmed or conversion_errors or debug"` -> `19 passed, 92 deselected`
+- `../.venv/bin/python -m pytest plugins/filters/async-context-compression/test_async_context_compression.py` -> `111 passed`
+- `git diff --check` -> passed
+
+## Findings
+
+### P2 -- Moderate
+
+| # | File | Issue | Reviewer(s) | Confidence | Route |
+|---|------|-------|-------------|------------|-------|
+| 1 | `plugins/filters/async-context-compression/async_context_compression.py:1583` | `_select_applicable_summary_snapshot()` 现在允许 `current_refs is None`,但 debug 日志分支仍执行 `current_refs[:...]`。当 idless body 只能通过 DB prefix 生成 refs、已有 best snapshot、后续低分 snapshot 进入 `debug_mode=True` 日志路径时,会抛 `TypeError`,把本应可选的 debug 诊断变成 inlet 请求失败。 | correctness, maintainability | 0.90 | `safe_auto -> review-fixer`, requires verification |
+| 2 | `plugins/filters/async-context-compression/async_context_compression.py:2313` | `_unfold_db_branch_for_body_ref_fallback()` 只捕获 `ImportError`。如果 `convert_output_to_messages()` 因 Open WebUI 版本漂移、异常 `output` payload 或转换 bug 抛出普通异常,DB fallback 不会 fail closed,而是让 inlet 整体失败。该路径是可选兼容 fallback,应拒绝复用而不是中断请求。 | maintainability | 0.78 | `safe_auto -> review-fixer`, requires verification |
+| 3 | `plugins/filters/async-context-compression/async_context_compression.py:2203` | 同长度 idless body 省略 DB assistant `output` 的匹配分支缺少测试覆盖。代码允许 body 缺少 `output`、但模型可见 role/content/tool_calls 与 DB 消息一致时复用 DB refs;现有 folded tool/reasoning 测试主要覆盖 unfolded DB path,不能证明 same-length output-omission 分支。 | testing | 0.78 | `manual -> downstream-resolver`, requires verification |
+| 4 | `plugins/filters/async-context-compression/async_context_compression.py:1734` | failed assistant 过滤只测试了 `history.currentId` 重建路径,未测试直接 `chat["messages"]` 路径。实现同时过滤 branch messages 和 direct messages,但当前测试只覆盖 history branch。 | testing | 0.74 | `manual -> downstream-resolver`, requires verification |
+
+### P3 -- Low
+
+| # | File | Issue | Reviewer(s) | Confidence | Route |
+|---|------|-------|-------------|------------|-------|
+| 5 | `plugins/filters/async-context-compression/async_context_compression.py:2286` | unfolded body 中 assistant `metadata.tool_outputs_trimmed` 的容忍分支缺少正反测试。现有测试覆盖 tool message 的 `metadata.is_trimmed` 和 reasoning exact-match,但没有证明 assistant 内容 mismatch 但带 `tool_outputs_trimmed=True` 时会安全复用,也没有缺少该 metadata 时拒绝复用的负例。 | testing | 0.68 | `advisory -> human` |
+
+## Positive Findings
+
+- 核心修复方向是保守的:无 id body 不会被直接当作可证明 refs 使用,而是先要求 body 与完整 DB branch、或去掉 terminal assistant 后的 user-tip branch 逐条匹配。
+- terminal assistant placeholder 只在 body 能证明匹配 user-tip branch 时才被忽略;snapshot 如果覆盖被忽略的 assistant,会在 live refs 校验中被拒绝。
+- body/db coverage map 已把 DB 覆盖坐标映射回 body 坐标,inlet 构造 tail 时使用 `_summary_snapshot_current_body_coverage_count()`,避免 folded output 展开后误删可见 tail。
+- mismatch、folded tool、folded reasoning、terminal assistant 正向路径、history branch failed assistant 过滤都有直接回归测试。
+
+## Coverage
+
+- 已运行:`../.venv/bin/python -m pytest plugins/filters/async-context-compression/test_async_context_compression.py` -> `106 passed`。
+- 未运行 backend 测试。本次可审的 `extensions` 子仓库最新提交只包含插件文件;用户描述中的 `backend/open_webui/utils/middleware.py` SSE `data: [DONE]` 修复与 `backend/open_webui/test/utils/test_llm_error_masking.py` 不在该提交 diff 中,因此本 review 未覆盖该 backend 修复。
+- 当前子仓库仍有未提交的无关修改:`docs/zh/future_plugin_development_roadmap_cn.md`,未纳入本次 review。
+
+## Verdict
+
+**Ready with fixes.**
+
+核心的 idless request / terminal assistant placeholder summary reuse 修复没有发现 P0/P1 阻塞问题;但建议在合并前处理两个 P2 代码问题:
+
+1. debug 日志路径改用 `current_refs_for_snapshot`,或在 `current_refs is None` 时跳过 common-prefix debug 计算,并补 `debug_mode=True + 多 snapshot + idless fallback` 回归测试。
+2. `convert_output_to_messages()` 的 fallback 展开异常应 fail closed,捕获普通异常并返回不兼容/未展开路径,同时补异常转换测试。
+
+测试覆盖缺口可以随后补齐,但同长度 output-omission 分支和 direct `chat["messages"]` failed assistant 过滤建议优先补上。
diff --git a/docs/plugins/filters/async-context-compression.md b/docs/plugins/filters/async-context-compression.md
index 96517cc..a537243 100644
--- a/docs/plugins/filters/async-context-compression.md
+++ b/docs/plugins/filters/async-context-compression.md
@@ -1,6 +1,6 @@
# Async Context Compression Filter
-| By [Fu-Jie](https://github.com/Fu-Jie) · v1.7.0 | [⭐ Star this repo](https://github.com/Fu-Jie/openwebui-extensions) |
+| By [Fu-Jie](https://github.com/Fu-Jie) · v1.7.1 | [⭐ Star this repo](https://github.com/Fu-Jie/openwebui-extensions) |
| :--- | ---: |
|  |  |  |  |  |  |  |
@@ -21,12 +21,14 @@ When the selection dialog opens, search for this plugin, check it, and continue.
> [!IMPORTANT]
> If the official OpenWebUI Community version is already installed, remove it first. After that, Batch Install Plugins can keep this plugin updated in future runs.
-## What's new in 1.7.0
+## What's new in 1.7.1
- **Branch-aware summary reuse**: Cached summaries are now validated against message ids and payload fingerprints before reuse, so summaries from sibling branches or edited content are rejected instead of being injected into the wrong branch.
- **Single-table branch summary storage**: Branch-valid summary rows are stored in `chat_summary`, while untrusted count-only summaries are regenerated.
- **Branch-aware referenced chats**: Referenced chats can now reuse the largest valid prefix summary plus the uncovered active-branch tail. If that mixed reference needs a new summary, the continuation summary is saved back to `chat_summary` for the referenced chat and reused later.
-- **Safer upgrade checks**: Legacy tables that still enforce one row per chat are rebuilt, and schema inspection failures leave existing tables untouched instead of running destructive DDL.
+- **Idless request summary reuse**: Request bodies without stable message ids can now be matched against the persisted DB active branch before rejecting cached summaries, preventing long chats from falling back to raw history unnecessarily.
+- **Terminal assistant placeholder handling**: When OpenWebUI has inserted the in-progress assistant placeholder into the DB branch but the model-visible request ends at the latest user message, the filter can ignore that terminal placeholder only after the body proves it matches the user-tip branch.
+- **Safer schema and DB fallback behavior**: Legacy tables that still enforce one row per chat are rebuilt safely, folded output conversion failures fail closed, and debug logging is safe for idless fallback.
## What's new in 1.6.4
@@ -212,6 +214,6 @@ If this plugin has been useful, a star on [OpenWebUI Extensions](https://github.
## Changelog
-See [`v1.7.0` Release Notes](https://github.com/Fu-Jie/openwebui-extensions/blob/main/plugins/filters/async-context-compression/v1.7.0.md) for the release-specific summary.
+See [`v1.7.1` Release Notes](https://github.com/Fu-Jie/openwebui-extensions/blob/main/plugins/filters/async-context-compression/v1.7.1.md) for the release-specific summary.
See the full history on GitHub: [OpenWebUI Extensions](https://github.com/Fu-Jie/openwebui-extensions)
diff --git a/docs/plugins/filters/async-context-compression.zh.md b/docs/plugins/filters/async-context-compression.zh.md
index 5da3dc9..b4ff51e 100644
--- a/docs/plugins/filters/async-context-compression.zh.md
+++ b/docs/plugins/filters/async-context-compression.zh.md
@@ -1,6 +1,6 @@
# 异步上下文压缩过滤器
-| 作者:[Fu-Jie](https://github.com/Fu-Jie) · v1.7.0 | [⭐ 点个 Star 支持项目](https://github.com/Fu-Jie/openwebui-extensions) |
+| 作者:[Fu-Jie](https://github.com/Fu-Jie) · v1.7.1 | [⭐ 点个 Star 支持项目](https://github.com/Fu-Jie/openwebui-extensions) |
| :--- | ---: |
|  |  |  |  |  |  |  |
@@ -23,12 +23,14 @@
> [!IMPORTANT]
> 如果你已经安装了 OpenWebUI 官方社区里的同名版本,请先删除旧版本,否则重新安装时可能报错。删除后,Batch Install Plugins 后续就可以继续负责更新这个插件。
-## 1.7.0 版本更新
+## 1.7.1 版本更新
- **分支感知摘要复用**:缓存摘要现在会先用 message id 和 payload fingerprint 验证覆盖范围。来自 sibling 分支或编辑前内容的旧摘要会被拒绝,不会注入到错误分支。
- **单表分支摘要存储**:branch-valid 摘要行统一写入 `chat_summary`,缺少覆盖元数据的 count-only 摘要会重新生成。
- **分支感知引用聊天**:引用聊天现在可以复用当前 active branch 上最大的有效前缀摘要,并拼接未覆盖的原文 tail。如果因此生成 continuation summary,会写回被引用聊天自己的 `chat_summary`,后续引用可直接复用。
-- **更安全的升级检查**:仍通过 `chat_id` unique 限制每个 chat 只能一行的旧表会被重建;如果无法反射检查现有 schema,会保留表不动并禁用摘要持久化,不会执行破坏性 DDL。
+- **修复无 id 请求的摘要复用**:当 request body 没有稳定 message id 时,插件会先用数据库里的 active branch 严格对齐校验,再决定是否复用已有摘要,避免长对话误退回原始全文。
+- **处理末尾 assistant 占位消息**:OpenWebUI 可能先把“正在生成中”的 assistant 占位消息写入数据库,但实际发给模型的 body 只到最新 user message。现在只有在 body 能证明匹配 user-tip 分支时,才会忽略这个末尾占位消息。
+- **更安全的 schema 和 DB fallback 行为**:仍通过 `chat_id` unique 限制每个 chat 只能一行的旧表会被安全重建;folded output 转换异常会 fail closed,idless fallback 的 debug 日志不会中断请求。
## 1.6.4 版本更新
@@ -247,6 +249,6 @@ flowchart TD
## 更新日志
-请查看 [`v1.7.0` 版本发布说明](https://github.com/Fu-Jie/openwebui-extensions/blob/main/plugins/filters/async-context-compression/v1.7.0_CN.md) 获取本次版本的独立发布摘要。
+请查看 [`v1.7.1` 版本发布说明](https://github.com/Fu-Jie/openwebui-extensions/blob/main/plugins/filters/async-context-compression/v1.7.1_CN.md) 获取本次版本的独立发布摘要。
完整历史请查看 GitHub 项目: [OpenWebUI Extensions](https://github.com/Fu-Jie/openwebui-extensions)
diff --git a/docs/plugins/filters/index.md b/docs/plugins/filters/index.md
index 839c9d7..38ec325 100644
--- a/docs/plugins/filters/index.md
+++ b/docs/plugins/filters/index.md
@@ -22,7 +22,7 @@ Filters act as middleware in the message pipeline:
Reduces token consumption in long conversations with tunable compression styles, safer summary fallbacks, and clearer failure visibility.
- **Version:** 1.7.0
+ **Version:** 1.7.1
[:octicons-arrow-right-24: Documentation](async-context-compression.md)
diff --git a/docs/plugins/filters/index.zh.md b/docs/plugins/filters/index.zh.md
index 0b08955..54f4a22 100644
--- a/docs/plugins/filters/index.zh.md
+++ b/docs/plugins/filters/index.zh.md
@@ -22,7 +22,7 @@ Filter 充当消息管线中的中间件:
通过可调压缩风格、更稳健的摘要回退和更清晰的失败提示,降低长对话的 token 消耗并保持连贯性。
- **版本:** 1.7.0
+ **版本:** 1.7.1
[:octicons-arrow-right-24: 查看文档](async-context-compression.zh.md)
diff --git a/plugins/filters/async-context-compression/README.md b/plugins/filters/async-context-compression/README.md
index e2698e6..80c29b5 100644
--- a/plugins/filters/async-context-compression/README.md
+++ b/plugins/filters/async-context-compression/README.md
@@ -1,6 +1,6 @@
# Async Context Compression Filter
-| By [Fu-Jie](https://github.com/Fu-Jie) · v1.7.0 | [⭐ Star this repo](https://github.com/Fu-Jie/openwebui-extensions) |
+| By [Fu-Jie](https://github.com/Fu-Jie) · v1.7.1 | [⭐ Star this repo](https://github.com/Fu-Jie/openwebui-extensions) |
| :--- | ---: |
|  |  |  |  |  |  |  |
@@ -28,12 +28,14 @@ When the selection dialog opens, search for this plugin, check it, and continue.
- **Protected-head tracking**: Summary rows remember how many leading messages were kept outside the summary. If the current `keep_first` policy no longer preserves those messages, the row is not reused as branch-valid coverage.
- **Safe upgrade behavior**: Legacy summaries without coverage metadata are not trusted as coverage. The first turn after upgrading may send more raw context until a branch-valid summary row is generated.
-## What's new in 1.7.0
+## What's new in 1.7.1
- **Branch-aware summary reuse**: Cached summaries are validated against ordered message refs and payload fingerprints before reuse, so sibling-branch or edited-history summaries are not injected into the wrong branch.
- **Single-table summary storage**: Branch-valid historical rows are stored directly in `chat_summary`; all branch summaries live in that table.
- **Branch-aware referenced chats**: Referenced chats can now reuse the largest valid prefix summary plus the uncovered active-branch tail. If that mixed reference needs a new summary, the continuation summary is saved back to `chat_summary` for the referenced chat and reused later.
-- **Safer schema upgrades**: Legacy count-only summaries are not trusted as coverage, and old one-row-per-chat schemas are rebuilt so summaries can be regenerated safely.
+- **Idless request summary reuse**: Request bodies without stable message ids can now be matched against the persisted DB active branch before rejecting cached summaries, preventing long chats from falling back to raw history unnecessarily.
+- **Terminal assistant placeholder handling**: When OpenWebUI has inserted the in-progress assistant placeholder into the DB branch but the model-visible request ends at the latest user message, the filter can ignore that terminal placeholder only after the body proves it matches the user-tip branch.
+- **Safer schema and DB fallback behavior**: Legacy count-only summaries are not trusted as coverage, old one-row-per-chat schemas are rebuilt safely, folded output conversion failures fail closed, and debug logging is safe for idless fallback.
## What's new in 1.6.4
@@ -181,6 +183,7 @@ flowchart TD
| `summary_model` | `None` | Model for summaries. Strongly recommended to set a fast, economical model (e.g., `gemini-2.5-flash`, `deepseek-v3`). Falls back to the current chat model when empty. |
| `summary_model_max_context` | `0` | Input context window used to fit summary requests. If `0`, falls back to `model_thresholds` or global `max_context_tokens`. |
| `max_summary_tokens` | `16384` | Maximum output length for the generated summary. This is not the summary-input context limit, and must be strictly less than 80% of the effective summary input window (`summary_model_max_context`, or its fallback from `model_thresholds` / `max_context_tokens`). The remaining window is reserved so the next compression can include the previous summary plus new messages; invalid settings raise an error instead of being auto-adjusted. |
+| `summary_llm_timeout_seconds` | `180.0` | Maximum time to wait for the summary LLM request before skipping summary generation for that turn. Set to `0` to disable the timeout. |
| `summary_temperature` | `0.1` | Randomness for summary generation. Lower is more deterministic. |
| `summary_fail_mode` | `silent` | Controls what happens when the summary LLM call fails. `silent` logs the error and skips summary generation for that turn; `raise` preserves the previous hard-failure behavior. |
| `compression_style` | `balanced` | Controls summary compactness. `aggressive` minimizes tokens, `balanced` keeps key context with moderate detail, and `faithful` preserves more nuance and reasoning context. |
@@ -210,6 +213,6 @@ If this plugin has been useful, a star on [OpenWebUI Extensions](https://github.
## Changelog
-See [`v1.7.0` Release Notes](https://github.com/Fu-Jie/openwebui-extensions/blob/main/plugins/filters/async-context-compression/v1.7.0.md) for the release-specific summary.
+See [`v1.7.1` Release Notes](https://github.com/Fu-Jie/openwebui-extensions/blob/main/plugins/filters/async-context-compression/v1.7.1.md) for the release-specific summary.
See the full history on GitHub: [OpenWebUI Extensions](https://github.com/Fu-Jie/openwebui-extensions)
diff --git a/plugins/filters/async-context-compression/README_CN.md b/plugins/filters/async-context-compression/README_CN.md
index c6ac814..2ea8792 100644
--- a/plugins/filters/async-context-compression/README_CN.md
+++ b/plugins/filters/async-context-compression/README_CN.md
@@ -1,6 +1,6 @@
# 异步上下文压缩过滤器
-| 作者:[Fu-Jie](https://github.com/Fu-Jie) · v1.7.0 | [⭐ 点个 Star 支持项目](https://github.com/Fu-Jie/openwebui-extensions) |
+| 作者:[Fu-Jie](https://github.com/Fu-Jie) · v1.7.1 | [⭐ 点个 Star 支持项目](https://github.com/Fu-Jie/openwebui-extensions) |
| :--- | ---: |
|  |  |  |  |  |  |  |
@@ -30,12 +30,14 @@
- **受保护头部追踪**:摘要行会记录有多少开头消息是在摘要之外按原文保留的。如果当前 `keep_first` 策略已经不再保留这些消息,该摘要行不会作为 branch-valid 覆盖范围复用。
- **安全升级行为**:没有覆盖范围元数据的 legacy summary 不再被当成可信覆盖。升级后的第一轮对话可能会发送更多原始上下文,直到生成 branch-valid 摘要行。
-## 1.7.0 版本更新
+## 1.7.1 版本更新
- **分支感知摘要复用**:缓存摘要会先用有序 message refs 和 payload fingerprints 校验,来自 sibling 分支或编辑前历史的摘要不会注入到错误分支。
- **单表摘要存储**:branch-valid 历史摘要行直接保存在 `chat_summary`,所有分支摘要都在这张表里。
- **分支感知引用聊天**:引用聊天现在可以复用当前 active branch 上最大的有效前缀摘要,并拼接未覆盖的原文 tail。如果因此生成 continuation summary,会写回被引用聊天自己的 `chat_summary`,后续引用可直接复用。
-- **更安全的 schema 升级**:legacy count-only 摘要不会被当成可信覆盖,旧的一 chat 一行 schema 会重建,后续重新生成安全摘要。
+- **修复无 id 请求的摘要复用**:当 request body 没有稳定 message id 时,插件会先用数据库里的 active branch 严格对齐校验,再决定是否复用已有摘要,避免长对话误退回原始全文。
+- **处理末尾 assistant 占位消息**:OpenWebUI 可能先把“正在生成中”的 assistant 占位消息写入数据库,但实际发给模型的 body 只到最新 user message。现在只有在 body 能证明匹配 user-tip 分支时,才会忽略这个末尾占位消息。
+- **更安全的 schema 和 DB fallback 行为**:legacy count-only 摘要不会被当成可信覆盖,旧的一 chat 一行 schema 会安全重建;folded output 转换异常会 fail closed,idless fallback 的 debug 日志不会中断请求。
## 1.6.5 版本更新
@@ -197,6 +199,7 @@ flowchart TD
| `summary_model` | `None` | 用于生成摘要的模型 ID。**强烈建议**配置快速、经济、上下文窗口大的模型(如 `gemini-2.5-flash`、`deepseek-v3`)。留空则尝试复用当前对话模型。 |
| `summary_model_max_context` | `0` | 摘要请求可使用的输入上下文窗口。如果为 0,则回退到 `model_thresholds` 或全局 `max_context_tokens`。 |
| `max_summary_tokens` | `16384` | 生成摘要时允许的最大输出 Token 数。它不是摘要输入窗口上限,并且必须严格小于有效摘要输入窗口(`summary_model_max_context`,或从 `model_thresholds` / `max_context_tokens` 回退得到的窗口)的 80%。剩余窗口用于下次压缩时同时容纳旧摘要和新消息;不满足要求会报错,不会自动调整。 |
+| `summary_llm_timeout_seconds` | `180.0` | 摘要 LLM 请求最长等待秒数;超过后本轮跳过摘要生成。设置为 `0` 可关闭超时。 |
| `summary_temperature` | `0.1` | 控制摘要生成的随机性,较低的值结果更稳定。 |
| `summary_fail_mode` | `silent` | 控制摘要 LLM 调用失败时的行为。`silent` 会记录错误并跳过本轮摘要;`raise` 会保留之前的硬抛错行为。 |
| `compression_style` | `balanced` | 控制摘要压缩风格。`aggressive` 更省 token,`balanced` 在紧凑和保真之间取中间值,`faithful` 会尽量保留更多细节、论证和上下文层次。 |
@@ -251,6 +254,6 @@ flowchart TD
## 更新日志
-请查看 [`v1.7.0` 版本发布说明](https://github.com/Fu-Jie/openwebui-extensions/blob/main/plugins/filters/async-context-compression/v1.7.0_CN.md) 获取本次版本的独立发布摘要。
+请查看 [`v1.7.1` 版本发布说明](https://github.com/Fu-Jie/openwebui-extensions/blob/main/plugins/filters/async-context-compression/v1.7.1_CN.md) 获取本次版本的独立发布摘要。
完整历史请查看 GitHub 项目: [OpenWebUI Extensions](https://github.com/Fu-Jie/openwebui-extensions)
diff --git a/plugins/filters/async-context-compression/async-context-compression-review-gpt-5.5.md b/plugins/filters/async-context-compression/async-context-compression-review-gpt-5.5.md
new file mode 100644
index 0000000..7905bdd
--- /dev/null
+++ b/plugins/filters/async-context-compression/async-context-compression-review-gpt-5.5.md
@@ -0,0 +1,50 @@
+## Code Review Results
+
+**Scope:** `9bebe32..HEAD` in `plugins/filters/async-context-compression` (2 files, 173 insertions, 2 deletions)
+**Intent:** Fix async context compression failures by settling stale background-summary status paths and timing out stalled summary LLM requests.
+**Mode:** report-only artifact requested by user; no code fixes applied.
+
+**Reviewers:** correctness, testing, maintainability, reliability, performance, kieran-python, project-standards, agent-native, learnings
+- reliability -- new background task status and timeout failure modes
+- performance -- async timeout/cancellation behavior
+- kieran-python -- Python async and exception semantics
+- project-standards -- no plugin-local `AGENTS.md` / `CLAUDE.md` found
+- agent-native -- no agent-facing surface changed
+- learnings -- no plugin-local `docs/solutions` store found
+
+### P2 -- Moderate
+
+| # | File | Issue | Reviewer | Confidence | Route |
+|---|------|-------|----------|------------|-------|
+| 1 | `plugins/filters/async-context-compression/async_context_compression.py:6306` | `_emit_summary_terminal_status()` awaits `__event_emitter__` directly. If the initial `done:false` status succeeds but the terminal status emit raises, `_generate_summary_async()` falls into the outer exception handler, which also awaits the same emitter unguarded at `:6824`. That can turn the intended stale-status recovery path into another unhandled background-task failure. Make terminal/error status emission best-effort and log emitter failures without re-raising. | reliability | 0.78 | `safe_auto -> review-fixer` |
+| 2 | `plugins/filters/async-context-compression/async_context_compression.py:6675` | The new save-failure terminal status branch is not directly tested. `test_summary_save_progress_matches_final_prompt_shrink` accidentally reaches a falsey save result because its mock returns `None`, but it only asserts that some status exists, not that the branch emits a final `done:true` summary-error status and skips success. Add a targeted `_save_summary == False` regression test. | testing, correctness, maintainability, kieran-python | 0.95 | `safe_auto -> review-fixer` |
+| 3 | `plugins/filters/async-context-compression/test_async_context_compression.py:4455` | Timeout coverage only exercises `_call_summary_llm()` directly with `summary_fail_mode = "raise"`. Production default is `silent`, where timeout returns an empty summary and `_generate_summary_async()` must settle the frontend status. Add a background-path test that times out the LLM call under default silent mode and asserts the final status is `done:true`. | testing, correctness | 0.90 | `safe_auto -> review-fixer` |
+
+### P3 -- Low
+
+| # | File | Issue | Reviewer | Confidence | Route |
+|---|------|-------|----------|------------|-------|
+| 4 | `plugins/filters/async-context-compression/async_context_compression.py:7333` | The `summary_llm_timeout_seconds = 0` branch is untested. A small test should prove that disabling the timeout allows a slow-but-eventually-successful summary request to complete instead of being treated as an immediate timeout. | testing | 0.85 | `safe_auto -> review-fixer` |
+| 5 | `plugins/filters/async-context-compression/README.md:176` | The new operator-facing `summary_llm_timeout_seconds` valve is not listed in the README configuration table. Add it to `README.md` and `README_CN.md` so operators know the default, tuning behavior, and that `0` disables the timeout. | project-standards | 0.75 | `safe_auto -> review-fixer` |
+
+### Residual Risks
+
+- `asyncio.wait_for()` bounds cooperative async stalls, but it cannot force-stop provider code that blocks the event loop or suppresses cancellation inside `generate_chat_completion`.
+- No integration test covers the actual Open WebUI provider transport cleanup path under timeout; current coverage is unit-level.
+
+### Coverage
+
+- Suppressed: 0 findings below threshold after synthesis.
+- Spawned reviewers completed: correctness, testing, maintainability, reliability, performance, kieran-python.
+- Manual coverage: project-standards, agent-native, learnings, adversarial failure scenarios.
+- Verification run locally: `pytest -q plugins/filters/async-context-compression/test_async_context_compression.py` -> `113 passed`.
+- Verification run locally: `git diff --check 9bebe32..HEAD` -> clean.
+- Main checkout dirty files under `/Users/nex/orca/workspaces/open-webui/Nautilus` were unrelated and excluded from this review.
+
+---
+
+> **Verdict:** Ready with fixes
+>
+> **Reasoning:** The core timeout/status changes are directionally correct and the existing suite passes, but the new failure-recovery paths need a small best-effort emitter hardening pass plus focused regression tests for save failure and default silent timeout behavior.
+>
+> **Fix order:** Harden terminal/error status emission -> add save-failure terminal status test -> add default silent timeout background-path test -> add timeout-disabled branch test and README entries.
diff --git a/plugins/filters/async-context-compression/async_context_compression.py b/plugins/filters/async-context-compression/async_context_compression.py
index ad7ac11..f9cd20c 100644
--- a/plugins/filters/async-context-compression/async_context_compression.py
+++ b/plugins/filters/async-context-compression/async_context_compression.py
@@ -5,7 +5,7 @@
author_url: https://github.com/Fu-Jie/openwebui-extensions
funding_url: https://github.com/open-webui
description: Reduces token consumption in long conversations while maintaining coherence through intelligent summarization and message compression.
-version: 1.7.0
+version: 1.7.1
openwebui_id: b1655bc8-6de9-4cad-8cb5-a6f7829a02ce
license: MIT
@@ -1376,11 +1376,21 @@ def _annotate_summary_snapshot_selection(
current_coverage_count: int,
current_coverage_refs: List[Dict[str, str]],
protected_head_count: int,
+ body_coverage_count: Optional[int] = None,
) -> Any:
"""Attach current-branch coverage metadata to the selected snapshot."""
setattr(snapshot, "_current_coverage_count", current_coverage_count)
setattr(snapshot, "_current_coverage_refs", current_coverage_refs)
setattr(snapshot, "_current_protected_head_count", protected_head_count)
+ setattr(
+ snapshot,
+ "_current_body_coverage_count",
+ (
+ current_coverage_count
+ if body_coverage_count is None
+ else max(0, int(body_coverage_count))
+ ),
+ )
return snapshot
def _summary_snapshot_current_coverage_count(self, snapshot: Any) -> int:
@@ -1393,6 +1403,16 @@ def _summary_snapshot_current_coverage_count(self, snapshot: Any) -> int:
or 0
)
+ def _summary_snapshot_current_body_coverage_count(self, snapshot: Any) -> int:
+ return int(
+ getattr(
+ snapshot,
+ "_current_body_coverage_count",
+ self._summary_snapshot_current_coverage_count(snapshot),
+ )
+ or 0
+ )
+
def _summary_snapshot_current_coverage_refs(
self, snapshot: Any
) -> Optional[List[Dict[str, str]]]:
@@ -1427,7 +1447,7 @@ def _select_applicable_summary_snapshot(
) -> Optional[Any]:
"""Choose the best snapshot that is safe for the current active branch."""
current_refs = self._current_branch_refs(messages)
- if current_refs is None:
+ if current_refs is None and require_full_coverage:
if self.valves.debug_mode:
logger.info(
"[Summary Snapshot] Current messages do not expose stable refs; "
@@ -1438,10 +1458,20 @@ def _select_applicable_summary_snapshot(
if require_full_coverage:
safe_boundary = len(current_refs)
elif max_coverage_count is not None:
- safe_boundary = min(len(current_refs), max(0, int(max_coverage_count)))
+ original_count = (
+ len(current_refs)
+ if current_refs is not None
+ else self._get_original_history_count(messages)
+ )
+ safe_boundary = min(original_count, max(0, int(max_coverage_count)))
else:
+ original_count = (
+ len(current_refs)
+ if current_refs is not None
+ else self._get_original_history_count(messages)
+ )
safe_boundary = min(
- len(current_refs),
+ original_count,
max(0, self._calculate_target_compressed_count(messages)),
)
effective_keep_first = (
@@ -1464,6 +1494,22 @@ def _select_applicable_summary_snapshot(
count = int(getattr(snapshot, "compressed_message_count", 0) or 0)
if count <= 0 or count != len(snapshot_refs):
continue
+
+ current_refs_for_snapshot = current_refs
+ if current_refs_for_snapshot is None:
+ current_refs_for_snapshot = self._message_refs_for_prefix(
+ messages,
+ count,
+ )
+ if current_refs_for_snapshot is None:
+ if self.valves.debug_mode:
+ logger.info(
+ "[Summary Snapshot] Rejecting snapshot because current "
+ f"messages do not expose stable refs through covered "
+ f"prefix count={count}."
+ )
+ continue
+
protected_head_count = self._parse_protected_head_count_json(
getattr(snapshot, "covered_message_refs_json", None)
)
@@ -1482,7 +1528,7 @@ def _select_applicable_summary_snapshot(
rejection_reason,
) = self._snapshot_coverage_for_current_branch(
snapshot_refs,
- current_refs,
+ current_refs_for_snapshot,
safe_boundary,
live_message_refs_by_id=live_message_refs_by_id,
)
@@ -1528,7 +1574,7 @@ def _select_applicable_summary_snapshot(
best_snapshot = self._annotate_summary_snapshot_selection(
snapshot,
matched_current_count,
- current_refs[:matched_current_count],
+ current_refs_for_snapshot[:matched_current_count],
protected_head_count,
)
continue
@@ -1536,7 +1582,9 @@ def _select_applicable_summary_snapshot(
if self.valves.debug_mode:
common = self._refs_common_prefix_length(
snapshot_refs,
- current_refs[: min(len(snapshot_refs), len(current_refs))],
+ current_refs_for_snapshot[
+ : min(len(snapshot_refs), len(current_refs_for_snapshot))
+ ],
)
logger.info(
"[Summary Snapshot] Skipping lower-ranked/stale snapshot: "
@@ -1640,6 +1688,23 @@ def _reconstruct_active_history_branch(
sortable_messages.sort(key=lambda item: (item[0], item[1]))
return [message for _, _, message in sortable_messages]
+ def _is_failed_assistant_message(self, message: Dict[str, Any]) -> bool:
+ """Mirror OpenWebUI middleware's failed-assistant filter."""
+ return (
+ isinstance(message, dict)
+ and message.get("role") == "assistant"
+ and "error" in message
+ )
+
+ def _filter_model_visible_history_messages(
+ self, messages: List[Dict[str, Any]]
+ ) -> List[Dict[str, Any]]:
+ return [
+ message
+ for message in messages
+ if not self._is_failed_assistant_message(message)
+ ]
+
async def _load_full_chat_messages(self, chat_id: str) -> List[Dict[str, Any]]:
"""Load the full persisted chat history for summary decisions when available."""
if not chat_id or Chats is None:
@@ -1664,11 +1729,15 @@ async def _load_full_chat_messages(self, chat_id: str) -> List[Dict[str, Any]]:
history_messages, current_id
)
if branch_messages:
- return branch_messages
+ return self._filter_model_visible_history_messages(
+ branch_messages
+ )
direct_messages = chat_payload.get("messages")
if isinstance(direct_messages, list) and direct_messages:
- return deepcopy(direct_messages)
+ return self._filter_model_visible_history_messages(
+ deepcopy(direct_messages)
+ )
return []
@@ -2133,6 +2202,220 @@ def _native_db_overlap_is_body_compatible(
return ref_index == len(body_refs)
+ def _body_message_matches_db_branch_message(
+ self,
+ body_message: Dict[str, Any],
+ db_message: Dict[str, Any],
+ ) -> bool:
+ """Compare a request-body message with its persisted active-branch peer."""
+ if not isinstance(body_message, dict) or not isinstance(db_message, dict):
+ return False
+ if self._is_summary_message(body_message) or self._is_summary_message(
+ db_message
+ ):
+ return False
+ if self._is_external_reference_message(
+ body_message
+ ) or self._is_external_reference_message(db_message):
+ return False
+
+ body_ref = self._message_ref(body_message)
+ db_ref = self._message_ref(db_message)
+ if db_ref is None:
+ return False
+ if body_ref is not None and body_ref["id"] != db_ref["id"]:
+ return False
+
+ if self._message_fingerprint_payload(
+ body_message
+ ) == self._message_fingerprint_payload(db_message):
+ return True
+
+ # OpenWebUI request bodies can omit folded assistant `output` while the
+ # DB history still carries it. The model-visible role/content/tool-call
+ # shape must still match before DB refs are trusted for an idless body.
+ if body_message.get("output") not in (None, "", []):
+ return False
+
+ return self._message_fingerprint_payload(
+ body_message,
+ include_output=False,
+ ) == self._message_fingerprint_payload(
+ db_message,
+ include_output=False,
+ )
+
+ def _body_message_matches_unfolded_db_message(
+ self,
+ body_message: Dict[str, Any],
+ unfolded_db_message: Dict[str, Any],
+ ) -> bool:
+ """Compare a body message against an unfolded DB message."""
+ if not isinstance(body_message, dict) or not isinstance(
+ unfolded_db_message, dict
+ ):
+ return False
+ if body_message.get("role") != unfolded_db_message.get("role"):
+ return False
+
+ body_ref = self._message_ref(body_message)
+ db_ref = self._message_ref(unfolded_db_message)
+ if body_ref is not None and db_ref is not None and body_ref["id"] != db_ref["id"]:
+ return False
+
+ body_tool_call_id = body_message.get("tool_call_id")
+ db_tool_call_id = unfolded_db_message.get("tool_call_id")
+ if isinstance(body_tool_call_id, str) or isinstance(db_tool_call_id, str):
+ if body_tool_call_id != db_tool_call_id:
+ return False
+
+ if self._message_fingerprint_payload(
+ body_message,
+ include_output=False,
+ ) == self._message_fingerprint_payload(
+ unfolded_db_message,
+ include_output=False,
+ ):
+ return True
+
+ metadata = body_message.get("metadata", {})
+ if not isinstance(metadata, dict):
+ metadata = {}
+
+ if body_message.get("role") == "tool" and metadata.get("is_trimmed"):
+ return True
+
+ if (
+ body_message.get("role") == "assistant"
+ and metadata.get("tool_outputs_trimmed")
+ ):
+ return True
+
+ return False
+
+ def _unfold_db_branch_for_body_ref_fallback(
+ self,
+ db_messages: List[Dict[str, Any]],
+ ) -> tuple[List[Dict[str, Any]], List[int]]:
+ """Unfold DB messages and map folded DB boundaries to unfolded indices."""
+ unfolded_messages: List[Dict[str, Any]] = []
+ db_to_body_boundaries = [0]
+
+ for message in db_messages:
+ before_count = len(unfolded_messages)
+
+ if (
+ isinstance(message, dict)
+ and message.get("role") == "assistant"
+ and isinstance(message.get("output"), list)
+ and message.get("output")
+ ):
+ expanded_messages = []
+ try:
+ from open_webui.utils.misc import convert_output_to_messages
+
+ expanded_messages = convert_output_to_messages(
+ message["output"], raw=True
+ )
+ except ImportError:
+ expanded_messages = []
+ except Exception as exc:
+ logger.debug(
+ "[Summary Snapshot] Failed to unfold DB assistant output "
+ "for idless body fallback; rejecting DB ref fallback: %s",
+ exc,
+ )
+ return [], []
+
+ # Match OpenWebUI middleware.process_messages_with_output(): the
+ # model-facing body uses convert_output_to_messages for every
+ # assistant `output`, including reasoning-only outputs.
+ if expanded_messages:
+ unfolded_messages.extend(deepcopy(expanded_messages))
+ else:
+ clean_message = {k: v for k, v in message.items() if k != "output"}
+ unfolded_messages.append(clean_message)
+ elif isinstance(message, dict):
+ clean_message = {k: v for k, v in message.items() if k != "output"}
+ unfolded_messages.append(clean_message)
+ else:
+ unfolded_messages.append(message)
+
+ if len(unfolded_messages) == before_count:
+ unfolded_messages.append(deepcopy(message))
+ db_to_body_boundaries.append(len(unfolded_messages))
+
+ self._normalize_native_tool_call_ids(unfolded_messages)
+ return unfolded_messages, db_to_body_boundaries
+
+ def _body_to_db_coverage_map_for_ref_fallback(
+ self,
+ body_messages: List[Dict[str, Any]],
+ db_messages: List[Dict[str, Any]],
+ ) -> Optional[List[int]]:
+ """Return DB-count -> body-boundary map when DB refs can stand in."""
+ if not body_messages or not db_messages:
+ return None
+
+ if len(body_messages) == len(db_messages) and all(
+ self._body_message_matches_db_branch_message(body_message, db_message)
+ for body_message, db_message in zip(body_messages, db_messages)
+ ):
+ return list(range(len(db_messages) + 1))
+
+ unfolded_messages, db_to_body_boundaries = (
+ self._unfold_db_branch_for_body_ref_fallback(db_messages)
+ )
+ if len(body_messages) != len(unfolded_messages):
+ return None
+
+ if not all(
+ self._body_message_matches_unfolded_db_message(
+ body_message,
+ unfolded_message,
+ )
+ for body_message, unfolded_message in zip(
+ body_messages,
+ unfolded_messages,
+ )
+ ):
+ return None
+
+ return db_to_body_boundaries
+
+ def _compatible_db_branch_for_body_ref_fallback(
+ self,
+ body_messages: List[Dict[str, Any]],
+ db_messages: List[Dict[str, Any]],
+ ) -> tuple[Optional[List[Dict[str, Any]]], Optional[List[int]], bool]:
+ """Find the persisted branch shape matching the model-visible body.
+
+ OpenWebUI inserts the new assistant placeholder before middleware
+ filters run, then middleware rebuilds the LLM payload only up to the
+ latest user message. When the plugin reads history.currentId it sees
+ that terminal assistant too; exclude it only if the request body proves
+ the user-tip branch is the matching source.
+ """
+ db_to_body_boundaries = self._body_to_db_coverage_map_for_ref_fallback(
+ body_messages,
+ db_messages,
+ )
+ if db_to_body_boundaries is not None:
+ return db_messages, db_to_body_boundaries, False
+
+ if not db_messages or db_messages[-1].get("role") != "assistant":
+ return None, None, False
+
+ user_tip_messages = db_messages[:-1]
+ db_to_body_boundaries = self._body_to_db_coverage_map_for_ref_fallback(
+ body_messages,
+ user_tip_messages,
+ )
+ if db_to_body_boundaries is None:
+ return None, None, False
+
+ return user_tip_messages, db_to_body_boundaries, True
+
async def _load_chat_history_live_refs(
self, chat_id: str
) -> Optional[Dict[str, Dict[str, str]]]:
@@ -2965,6 +3248,11 @@ class Valves(BaseModel):
ge=1,
description="The maximum number of tokens for the summary. Must be less than 80% of the summary model input window so at least 20% remains for new messages during follow-up compression.",
)
+ summary_llm_timeout_seconds: float = Field(
+ default=180.0,
+ ge=0,
+ description="Maximum seconds to wait for the summary LLM request. Set to 0 to disable the timeout.",
+ )
def __init__(self, **data):
super().__init__(**data)
@@ -3821,7 +4109,7 @@ async def _load_applicable_summary_snapshot(
return None
if live_message_refs_by_id is None:
live_message_refs_by_id = await self._load_chat_history_live_refs(chat_id)
- return self._select_applicable_summary_snapshot(
+ selected = self._select_applicable_summary_snapshot(
snapshots,
messages,
require_full_coverage=require_full_coverage,
@@ -3829,6 +4117,58 @@ async def _load_applicable_summary_snapshot(
max_coverage_count=max_coverage_count,
enforce_keep_first=enforce_keep_first,
)
+ if selected is not None:
+ return selected
+
+ if require_full_coverage or self._current_branch_refs(messages) is not None:
+ return None
+
+ db_messages = await self._load_full_chat_messages(chat_id)
+ (
+ compatible_db_messages,
+ db_to_body_boundaries,
+ ignored_terminal_assistant,
+ ) = self._compatible_db_branch_for_body_ref_fallback(
+ messages,
+ db_messages,
+ )
+ if compatible_db_messages is None or db_to_body_boundaries is None:
+ if self.valves.debug_mode:
+ logger.info(
+ "[Summary Snapshot] DB active branch fallback skipped: "
+ f"request body is not a compatible idless view "
+ f"(body_count={len(messages)}, db_count={len(db_messages)})."
+ )
+ return None
+
+ selected = self._select_applicable_summary_snapshot(
+ snapshots,
+ compatible_db_messages,
+ require_full_coverage=require_full_coverage,
+ live_message_refs_by_id=live_message_refs_by_id,
+ max_coverage_count=max_coverage_count,
+ enforce_keep_first=enforce_keep_first,
+ )
+ if selected is not None:
+ db_coverage_count = self._summary_snapshot_current_coverage_count(selected)
+ if db_coverage_count >= len(db_to_body_boundaries):
+ return None
+ self._annotate_summary_snapshot_selection(
+ selected,
+ db_coverage_count,
+ self._summary_snapshot_current_coverage_refs(selected) or [],
+ self._summary_snapshot_current_protected_head_count(selected),
+ body_coverage_count=db_to_body_boundaries[db_coverage_count],
+ )
+ if self.valves.debug_mode:
+ logger.info(
+ "[Summary Snapshot] Selected snapshot using DB active branch refs "
+ "for request body without stable message ids "
+ f"(db_coverage={db_coverage_count}, "
+ f"body_coverage={db_to_body_boundaries[db_coverage_count]}, "
+ f"ignored_terminal_assistant={ignored_terminal_assistant})."
+ )
+ return selected
def _count_tokens(self, text: str) -> int:
"""Counts the number of tokens in the text."""
@@ -5022,6 +5362,9 @@ async def inlet(
compressed_count = self._summary_snapshot_current_coverage_count(
summary_snapshot
)
+ body_compressed_count = (
+ self._summary_snapshot_current_body_coverage_count(summary_snapshot)
+ )
covered_refs = self._summary_snapshot_current_coverage_refs(
summary_snapshot
)
@@ -5038,7 +5381,7 @@ async def inlet(
# 2. Tail messages (Tail) - All messages starting from the last compression point.
# Align legacy/raw progress to an atomic boundary so old summary rows do not
# reintroduce orphaned tool messages into the retained tail.
- raw_start_index = max(compressed_count, effective_keep_first)
+ raw_start_index = max(body_compressed_count, effective_keep_first)
start_index = self._align_tail_start_to_atomic_boundary(
messages, raw_start_index, effective_keep_first
)
@@ -5858,14 +6201,16 @@ async def _check_and_generate_summary_async(
lang, "status_high_usage"
)
- await __event_emitter__(
+ await self._emit_status_event(
+ __event_emitter__,
{
"type": "status",
"data": {
"description": status_msg,
"done": True,
},
- }
+ },
+ "[🔍 Background Calculation] context usage status",
)
# Check if compression is needed
@@ -5930,18 +6275,19 @@ async def _check_and_generate_summary_async(
event_call=__event_call__,
force=True,
)
- if __event_emitter__:
- await __event_emitter__(
- {
- "type": "status",
- "data": {
- "description": self._get_translation(
- lang, "status_summary_error", error=str(e)[:100]
- ),
- "done": True,
- },
- }
- )
+ await self._emit_status_event(
+ __event_emitter__,
+ {
+ "type": "status",
+ "data": {
+ "description": self._get_translation(
+ lang, "status_summary_error", error=str(e)[:100]
+ ),
+ "done": True,
+ },
+ },
+ "[🔍 Background Calculation] error status",
+ )
logger.exception("[🔍 Background Calculation] Unhandled exception")
def _clean_model_id(self, model_id: Optional[str]) -> Optional[str]:
@@ -5951,6 +6297,47 @@ def _clean_model_id(self, model_id: Optional[str]) -> Optional[str]:
cleaned = model_id.strip().strip('"').strip("'")
return cleaned if cleaned else None
+ async def _emit_status_event(
+ self,
+ __event_emitter__: Optional[Callable[[Any], Awaitable[None]]],
+ event: Dict[str, Any],
+ context: str,
+ ) -> bool:
+ if not __event_emitter__:
+ return False
+
+ try:
+ await __event_emitter__(event)
+ return True
+ except Exception as exc:
+ logger.warning(
+ "%s emission failed: %s: %s",
+ context,
+ type(exc).__name__,
+ exc,
+ )
+ return False
+
+ async def _emit_summary_terminal_status(
+ self,
+ __event_emitter__: Callable[[Any], Awaitable[None]],
+ lang: str,
+ reason: str,
+ ) -> None:
+ await self._emit_status_event(
+ __event_emitter__,
+ {
+ "type": "status",
+ "data": {
+ "description": self._get_translation(
+ lang, "status_summary_error", error=str(reason)[:100]
+ ),
+ "done": True,
+ },
+ },
+ "[🤖 Async Summary Task] terminal status",
+ )
+
async def _generate_summary_async(
self,
messages: list,
@@ -6212,18 +6599,19 @@ async def _generate_summary_async(
# 6. Call LLM to generate new summary
# Send status notification for starting summary generation
- if __event_emitter__:
- await __event_emitter__(
- {
- "type": "status",
- "data": {
- "description": self._get_translation(
- lang, "status_generating_summary"
- ),
- "done": False,
- },
- }
- )
+ await self._emit_status_event(
+ __event_emitter__,
+ {
+ "type": "status",
+ "data": {
+ "description": self._get_translation(
+ lang, "status_generating_summary"
+ ),
+ "done": False,
+ },
+ },
+ "[🤖 Async Summary Task] generating status",
+ )
new_summary = await self._call_summary_llm(
conversation_text,
@@ -6240,6 +6628,11 @@ async def _generate_summary_async(
log_type="warning",
event_call=__event_call__,
)
+ await self._emit_summary_terminal_status(
+ __event_emitter__,
+ lang,
+ "summary generation returned empty result",
+ )
return
if summary_index is None:
@@ -6303,23 +6696,29 @@ async def _generate_summary_async(
log_type="warning",
event_call=__event_call__,
)
+ await self._emit_summary_terminal_status(
+ __event_emitter__,
+ lang,
+ "summary generated but was not persisted",
+ )
return
# Send completion status notification
- if __event_emitter__:
- await __event_emitter__(
- {
- "type": "status",
- "data": {
- "description": self._get_translation(
- lang,
- "status_loaded_summary",
- count=len(middle_messages),
- ),
- "done": True,
- },
- }
- )
+ await self._emit_status_event(
+ __event_emitter__,
+ {
+ "type": "status",
+ "data": {
+ "description": self._get_translation(
+ lang,
+ "status_loaded_summary",
+ count=len(middle_messages),
+ ),
+ "done": True,
+ },
+ },
+ "[🤖 Async Summary Task] completion status",
+ )
await self._log(
f"[🤖 Async Summary Task] ✅ Complete! New summary length: {len(new_summary)} characters",
@@ -6446,18 +6845,19 @@ async def _generate_summary_async(
force=True,
)
- if __event_emitter__:
- await __event_emitter__(
- {
- "type": "status",
- "data": {
- "description": self._get_translation(
- lang, "status_summary_error", error=str(e)[:100]
- ),
- "done": True,
- },
- }
- )
+ await self._emit_status_event(
+ __event_emitter__,
+ {
+ "type": "status",
+ "data": {
+ "description": self._get_translation(
+ lang, "status_summary_error", error=str(e)[:100]
+ ),
+ "done": True,
+ },
+ },
+ "[🤖 Async Summary Task] error status",
+ )
import traceback
@@ -6956,7 +7356,14 @@ async def _call_summary_llm(
request = __request__ or Request(scope={"type": "http", "app": webui_app})
# Call generate_chat_completion
- response = await generate_chat_completion(request, payload, user)
+ summary_timeout = float(self.valves.summary_llm_timeout_seconds or 0)
+ if summary_timeout > 0:
+ response = await asyncio.wait_for(
+ generate_chat_completion(request, payload, user),
+ timeout=summary_timeout,
+ )
+ else:
+ response = await generate_chat_completion(request, payload, user)
# Handle JSONResponse (some backends return JSONResponse instead of dict)
if hasattr(response, "body"):
@@ -7007,7 +7414,12 @@ async def _call_summary_llm(
except Exception as e:
error_msg = str(e)
# Handle specific error messages
- if "Model not found" in error_msg:
+ if isinstance(e, asyncio.TimeoutError):
+ timeout_display = f"{float(self.valves.summary_llm_timeout_seconds):g}"
+ error_message = (
+ f"Summary LLM request timed out after {timeout_display} seconds."
+ )
+ elif "Model not found" in error_msg:
error_message = f"Summary model '{model}' not found."
else:
error_message = f"Summary LLM Error ({model}): {error_msg}"
diff --git a/plugins/filters/async-context-compression/test_async_context_compression.py b/plugins/filters/async-context-compression/test_async_context_compression.py
index 0d617fe..da8b01b 100644
--- a/plugins/filters/async-context-compression/test_async_context_compression.py
+++ b/plugins/filters/async-context-compression/test_async_context_compression.py
@@ -103,6 +103,7 @@ def _install_openwebui_stubs() -> None:
_ensure_module("open_webui")
_ensure_module("open_webui.utils")
chat_module = _ensure_module("open_webui.utils.chat")
+ misc_module = _ensure_module("open_webui.utils.misc")
_ensure_module("open_webui.models")
users_module = _ensure_module("open_webui.models.users")
models_module = _ensure_module("open_webui.models.models")
@@ -114,6 +115,9 @@ def _install_openwebui_stubs() -> None:
async def generate_chat_completion(*args, **kwargs):
return {}
+ def convert_output_to_messages(output, raw=False):
+ return deepcopy(output) if isinstance(output, list) else []
+
class DummyUsers:
pass
@@ -132,6 +136,7 @@ def __init__(self, *args, **kwargs):
pass
chat_module.generate_chat_completion = generate_chat_completion
+ misc_module.convert_output_to_messages = convert_output_to_messages
users_module.Users = DummyUsers
models_module.Models = DummyModels
chats_module.Chats = DummyChats
@@ -414,6 +419,12 @@ async def fake_summary_llm(
class TestAsyncContextCompression(unittest.TestCase):
def setUp(self):
+ misc_module = _ensure_module("open_webui.utils.misc")
+ misc_module.convert_output_to_messages = (
+ lambda output, raw=False: deepcopy(output)
+ if isinstance(output, list)
+ else []
+ )
self.filter = module.Filter()
def test_build_summary_prompt_defaults_to_balanced_style(self):
@@ -1450,15 +1461,20 @@ async def noop(*args, **kwargs):
["m0", "m1", "m3", "m4"],
)
- def test_outlet_does_not_reinject_live_sibling_snapshot(self):
+ def test_inlet_reuses_prefix_snapshot_when_later_tool_tail_lacks_ids(self):
self.filter.valves.keep_last = 0
- current_messages = _messages_with_ids(["m0", "m1", "new_m2", "new_m3"])
- old_branch_messages = _messages_with_ids(["m0", "m1", "old_m2", "old_m3"])
- old_refs = self.filter._message_refs_for_prefix(old_branch_messages, 4)
- snapshots = [_snapshot("old branch summary", old_refs)]
- live_messages = current_messages + old_branch_messages[2:]
- captured = {}
- scheduled = []
+ stable_prefix = _messages_with_ids(["m0", "m1", "m2", "m3"])
+ current_messages = stable_prefix + [
+ {
+ "role": "assistant",
+ "content": "calling tool",
+ "tool_calls": [{"id": "call_1", "type": "function"}],
+ },
+ {"role": "tool", "tool_call_id": "call_1", "content": "tool result"},
+ {"role": "user", "content": "new question"},
+ ]
+ prefix_refs = self.filter._message_refs_for_prefix(stable_prefix, 4)
+ snapshots = [_snapshot("summary before idless tool tail", prefix_refs)]
async def fake_load_snapshot(
chat_id,
@@ -1469,186 +1485,837 @@ async def fake_load_snapshot(
snapshots,
messages,
require_full_coverage=require_full_coverage,
- live_message_refs_by_id=_live_refs_by_id(self.filter, live_messages),
+ live_message_refs_by_id=_live_refs_by_id(self.filter, stable_prefix),
)
- async def fake_user_context(__user__, __event_call__):
- return {"user_language": "en-US"}
-
- async def fake_locked_summary_task(
- lock,
- chat_id,
- model,
- body,
- user_data,
- target_compressed_count,
- lang,
- __event_emitter__,
- __event_call__,
- __request__=None,
- ):
- captured["messages"] = body["messages"]
-
async def noop(*args, **kwargs):
return None
- def fake_create_task(coro):
- scheduled.append(coro)
- return None
-
self.filter._load_applicable_summary_snapshot = fake_load_snapshot
- self.filter._get_user_context = fake_user_context
- self.filter._get_chat_context = lambda body, metadata=None: {
- "chat_id": "chat-1",
- "message_id": "msg-1",
- }
- self.filter._should_skip_compression = lambda body, model: False
- self.filter._locked_summary_task = fake_locked_summary_task
self.filter._log = noop
+ self.filter._emit_debug_log = noop
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 0
+ }
- original_create_task = asyncio.create_task
- asyncio.create_task = fake_create_task
- try:
- asyncio.run(
- self.filter.outlet(
- {"model": "test-model", "messages": current_messages},
- __event_call__=None,
- )
- )
- finally:
- asyncio.create_task = original_create_task
+ body = {
+ "chat_id": "chat-1",
+ "model": "test-model",
+ "messages": current_messages,
+ }
- self.assertEqual(len(scheduled), 1)
- asyncio.run(scheduled[0])
+ result = asyncio.run(self.filter.inlet(body))
+ final_messages = result["messages"]
- self.assertFalse(
- any(self.filter._is_summary_message(message) for message in captured["messages"])
- )
+ self.assertEqual(len(final_messages), 4)
+ self.assertTrue(self.filter._is_summary_message(final_messages[0]))
+ self.assertIn("summary before idless tool tail", final_messages[0]["content"])
+ self.assertEqual(final_messages[1]["content"], "calling tool")
+ self.assertEqual(final_messages[2]["role"], "tool")
+ self.assertEqual(final_messages[3]["content"], "new question")
self.assertEqual(
- [message["id"] for message in captured["messages"]],
- ["m0", "m1", "new_m2", "new_m3"],
+ [
+ ref["id"]
+ for ref in final_messages[0]["metadata"]["covered_message_refs"]
+ ],
+ ["m0", "m1", "m2", "m3"],
)
- def test_outlet_reinjects_matching_branch_snapshot_with_metadata(self):
+ def test_snapshot_selection_rejects_snapshot_reaching_idless_tool_tail(self):
self.filter.valves.keep_last = 0
- current_messages = _messages_with_ids(["m0", "m1", "m2", "m3", "m4"])
- prefix_refs = self.filter._message_refs_for_prefix(current_messages, 3)
- snapshots = [_snapshot("shared prefix summary", prefix_refs)]
- captured = {}
- scheduled = []
+ stable_prefix = _messages_with_ids(["m0", "m1", "m2", "m3"])
+ current_messages = stable_prefix + [
+ {
+ "role": "assistant",
+ "content": "calling tool",
+ "tool_calls": [{"id": "call_1", "type": "function"}],
+ },
+ {"role": "tool", "tool_call_id": "call_1", "content": "tool result"},
+ ]
+ unsafe_refs = self.filter._message_refs_for_prefix(
+ stable_prefix
+ + [
+ {
+ "id": "tool-call",
+ "role": "assistant",
+ "content": "calling tool",
+ "tool_calls": [{"id": "call_1", "type": "function"}],
+ }
+ ],
+ 5,
+ )
- async def fake_load_snapshot(
- chat_id,
- messages,
- require_full_coverage=False,
- ):
- return self.filter._select_applicable_summary_snapshot(
- snapshots,
- messages,
- require_full_coverage=require_full_coverage,
- live_message_refs_by_id=_live_refs_by_id(self.filter, current_messages),
+ selected = self.filter._select_applicable_summary_snapshot(
+ [_snapshot("unsafe tool summary", unsafe_refs)],
+ current_messages,
+ live_message_refs_by_id=_live_refs_by_id(self.filter, stable_prefix),
+ )
+
+ self.assertIsNone(selected)
+
+ def test_inlet_reuses_db_branch_snapshot_when_body_has_no_message_ids(self):
+ self.filter.valves.keep_last = 0
+ db_messages = _messages_with_ids([f"m{i}" for i in range(7)])
+ body_messages = [
+ {
+ key: deepcopy(value)
+ for key, value in message.items()
+ if key != "id"
+ }
+ for message in db_messages
+ ]
+ snapshots = [
+ _snapshot(
+ "db-backed summary",
+ self.filter._message_refs_for_prefix(db_messages, 4),
)
+ ]
- async def fake_user_context(__user__, __event_call__):
- return {"user_language": "en-US"}
+ async def fake_load_snapshots(chat_id):
+ return snapshots
- async def fake_locked_summary_task(
- lock,
- chat_id,
- model,
- body,
- user_data,
- target_compressed_count,
- lang,
- __event_emitter__,
- __event_call__,
- __request__=None,
- ):
- captured["messages"] = body["messages"]
+ async def fake_load_live_refs(chat_id):
+ return _live_refs_by_id(self.filter, db_messages)
- async def noop(*args, **kwargs):
- return None
+ async def fake_load_full_chat_messages(chat_id):
+ return db_messages
- def fake_create_task(coro):
- scheduled.append(coro)
+ async def noop(*args, **kwargs):
return None
- self.filter._load_applicable_summary_snapshot = fake_load_snapshot
- self.filter._get_user_context = fake_user_context
- self.filter._get_chat_context = lambda body, metadata=None: {
- "chat_id": "chat-1",
- "message_id": "msg-1",
- }
- self.filter._should_skip_compression = lambda body, model: False
- self.filter._locked_summary_task = fake_locked_summary_task
+ self.filter._load_summary_snapshots = fake_load_snapshots
+ self.filter._load_chat_history_live_refs = fake_load_live_refs
+ self.filter._load_full_chat_messages = fake_load_full_chat_messages
self.filter._log = noop
+ self.filter._emit_debug_log = noop
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 0
+ }
- original_create_task = asyncio.create_task
- asyncio.create_task = fake_create_task
- try:
- asyncio.run(
- self.filter.outlet(
- {"model": "test-model", "messages": current_messages},
- __event_call__=None,
- )
+ result = asyncio.run(
+ self.filter.inlet(
+ {
+ "chat_id": "chat-1",
+ "model": "test-model",
+ "messages": body_messages,
+ }
)
- finally:
- asyncio.create_task = original_create_task
-
- self.assertEqual(len(scheduled), 1)
- asyncio.run(scheduled[0])
- final_messages = captured["messages"]
+ )
+ final_messages = result["messages"]
- self.assertEqual(len(final_messages), 6)
- self.assertTrue(self.filter._is_summary_message(final_messages[3]))
+ self.assertEqual(len(final_messages), 4)
+ self.assertTrue(self.filter._is_summary_message(final_messages[0]))
+ self.assertIn("db-backed summary", final_messages[0]["content"])
+ self.assertEqual(
+ [message["content"] for message in final_messages[1:]],
+ ["message m4", "message m5", "message m6"],
+ )
self.assertEqual(
[
ref["id"]
- for ref in final_messages[3]["metadata"]["covered_message_refs"]
+ for ref in final_messages[0]["metadata"]["covered_message_refs"]
],
- ["m0", "m1", "m2"],
+ ["m0", "m1", "m2", "m3"],
)
- self.assertEqual(final_messages[4]["id"], "m3")
- self.assertEqual(final_messages[5]["id"], "m4")
- def test_snapshot_selection_rejects_same_content_different_ids(self):
+ def test_inlet_reuses_db_snapshot_when_current_tip_is_assistant_placeholder(self):
self.filter.valves.keep_last = 0
- current_messages = [
- {"id": "new-1", "role": "user", "content": "same"},
- {"id": "new-2", "role": "assistant", "content": "same"},
+ db_messages = _messages_with_ids([f"m{i}" for i in range(6)])
+ self.assertEqual(db_messages[-2]["role"], "user")
+ self.assertEqual(db_messages[-1]["role"], "assistant")
+ body_messages = [
+ {
+ key: deepcopy(value)
+ for key, value in message.items()
+ if key != "id"
+ }
+ for message in db_messages[:-1]
]
- old_messages = [
- {"id": "old-1", "role": "user", "content": "same"},
- {"id": "old-2", "role": "assistant", "content": "same"},
+ snapshots = [
+ _snapshot(
+ "user-tip db summary",
+ self.filter._message_refs_for_prefix(db_messages, 4),
+ )
]
- old_refs = self.filter._message_refs_for_prefix(old_messages, 2)
- selected = self.filter._select_applicable_summary_snapshot(
- [_snapshot("old same content", old_refs)],
- current_messages,
- live_message_refs_by_id=_live_refs_by_id(
- self.filter, current_messages + old_messages
- ),
+ async def fake_load_snapshots(chat_id):
+ return snapshots
+
+ async def fake_load_live_refs(chat_id):
+ return _live_refs_by_id(self.filter, db_messages)
+
+ async def fake_load_full_chat_messages(chat_id):
+ return db_messages
+
+ async def noop(*args, **kwargs):
+ return None
+
+ self.filter._load_summary_snapshots = fake_load_snapshots
+ self.filter._load_chat_history_live_refs = fake_load_live_refs
+ self.filter._load_full_chat_messages = fake_load_full_chat_messages
+ self.filter._log = noop
+ self.filter._emit_debug_log = noop
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 0
+ }
+
+ result = asyncio.run(
+ self.filter.inlet(
+ {
+ "chat_id": "chat-1",
+ "model": "test-model",
+ "messages": body_messages,
+ }
+ )
)
+ final_messages = result["messages"]
- self.assertIsNone(selected)
+ self.assertEqual(len(final_messages), 2)
+ self.assertTrue(self.filter._is_summary_message(final_messages[0]))
+ self.assertIn("user-tip db summary", final_messages[0]["content"])
+ self.assertEqual(final_messages[1]["content"], "message m4")
+ self.assertNotIn("message m5", [message["content"] for message in final_messages])
+ self.assertEqual(
+ [
+ ref["id"]
+ for ref in final_messages[0]["metadata"]["covered_message_refs"]
+ ],
+ ["m0", "m1", "m2", "m3"],
+ )
- def test_snapshot_selection_rejects_same_id_changed_payload(self):
+ def test_snapshot_selection_debug_handles_idless_tail_with_multiple_snapshots(self):
self.filter.valves.keep_last = 0
- original_messages = [
- {"id": "m1", "role": "user", "content": "original question"},
- {"id": "m2", "role": "assistant", "content": "original answer"},
- ]
- edited_messages = [
- {"id": "m1", "role": "user", "content": "edited question"},
- {"id": "m2", "role": "assistant", "content": "original answer"},
+ self.filter.valves.debug_mode = True
+ messages = _messages_with_ids([f"m{i}" for i in range(5)]) + [
+ {"role": "assistant", "content": "idless visible tail"}
]
- original_refs = self.filter._message_refs_for_prefix(original_messages, 2)
+ larger_snapshot = _snapshot(
+ "larger summary",
+ self.filter._message_refs_for_prefix(messages, 4),
+ )
+ smaller_snapshot = _snapshot(
+ "smaller summary",
+ self.filter._message_refs_for_prefix(messages, 3),
+ )
selected = self.filter._select_applicable_summary_snapshot(
- [_snapshot("old edited content", original_refs)],
- edited_messages,
- live_message_refs_by_id=_live_refs_by_id(self.filter, edited_messages),
+ [larger_snapshot, smaller_snapshot],
+ messages,
+ live_message_refs_by_id=_live_refs_by_id(self.filter, messages[:5]),
+ )
+
+ self.assertIs(selected, larger_snapshot)
+
+ def test_db_branch_snapshot_fallback_rejects_mismatched_idless_body(self):
+ self.filter.valves.keep_last = 0
+ db_messages = _messages_with_ids([f"m{i}" for i in range(5)])
+ body_messages = [
+ {
+ key: deepcopy(value)
+ for key, value in message.items()
+ if key != "id"
+ }
+ for message in db_messages
+ ]
+ body_messages[2]["content"] = "edited body payload"
+ snapshots = [
+ _snapshot(
+ "unsafe db summary",
+ self.filter._message_refs_for_prefix(db_messages, 3),
+ )
+ ]
+
+ async def fake_load_snapshots(chat_id):
+ return snapshots
+
+ async def fake_load_live_refs(chat_id):
+ return _live_refs_by_id(self.filter, db_messages)
+
+ async def fake_load_full_chat_messages(chat_id):
+ return db_messages
+
+ async def noop(*args, **kwargs):
+ return None
+
+ self.filter._load_summary_snapshots = fake_load_snapshots
+ self.filter._load_chat_history_live_refs = fake_load_live_refs
+ self.filter._load_full_chat_messages = fake_load_full_chat_messages
+ self.filter._log = noop
+ self.filter._emit_debug_log = noop
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 0
+ }
+
+ result = asyncio.run(
+ self.filter.inlet(
+ {
+ "chat_id": "chat-1",
+ "model": "test-model",
+ "messages": body_messages,
+ }
+ )
+ )
+
+ self.assertFalse(
+ any(
+ self.filter._is_summary_message(message)
+ for message in result["messages"]
+ )
+ )
+ self.assertEqual(result["messages"], body_messages)
+
+ def test_inlet_reuses_same_length_idless_body_that_omits_db_output(self):
+ self.filter.valves.keep_last = 0
+ db_messages = [
+ {"id": "m0", "role": "user", "content": "message m0"},
+ {
+ "id": "m1",
+ "role": "assistant",
+ "content": "visible answer",
+ "tool_calls": [
+ {
+ "id": "call_1",
+ "type": "function",
+ "function": {"name": "search", "arguments": "{}"},
+ }
+ ],
+ "output": [{"role": "assistant", "content": "hidden folded output"}],
+ },
+ {"id": "m2", "role": "user", "content": "message m2"},
+ ]
+ body_messages = [
+ {"role": "user", "content": "message m0"},
+ {
+ "role": "assistant",
+ "content": "visible answer",
+ "tool_calls": deepcopy(db_messages[1]["tool_calls"]),
+ },
+ {"role": "user", "content": "message m2"},
+ ]
+ snapshots = [
+ _snapshot(
+ "same length output omitted summary",
+ self.filter._message_refs_for_prefix(db_messages, 2),
+ )
+ ]
+
+ async def fake_load_snapshots(chat_id):
+ return snapshots
+
+ async def fake_load_live_refs(chat_id):
+ return _live_refs_by_id(self.filter, db_messages)
+
+ async def fake_load_full_chat_messages(chat_id):
+ return db_messages
+
+ async def noop(*args, **kwargs):
+ return None
+
+ self.filter._load_summary_snapshots = fake_load_snapshots
+ self.filter._load_chat_history_live_refs = fake_load_live_refs
+ self.filter._load_full_chat_messages = fake_load_full_chat_messages
+ self.filter._log = noop
+ self.filter._emit_debug_log = noop
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 0
+ }
+
+ result = asyncio.run(
+ self.filter.inlet(
+ {
+ "chat_id": "chat-1",
+ "model": "test-model",
+ "messages": body_messages,
+ }
+ )
+ )
+ final_messages = result["messages"]
+
+ self.assertTrue(self.filter._is_summary_message(final_messages[0]))
+ self.assertIn("same length output omitted summary", final_messages[0]["content"])
+ self.assertEqual(final_messages[1]["content"], "message m2")
+
+ mismatched_body = deepcopy(body_messages)
+ mismatched_body[1]["tool_calls"][0]["function"]["name"] = "other"
+ self.assertIsNone(
+ self.filter._body_to_db_coverage_map_for_ref_fallback(
+ mismatched_body,
+ db_messages,
+ )
+ )
+
+ def test_unfold_db_branch_fallback_rejects_conversion_errors(self):
+ misc_module = _ensure_module("open_webui.utils.misc")
+
+ def convert_output_to_messages(output, raw=False):
+ raise RuntimeError("bad folded output")
+
+ misc_module.convert_output_to_messages = convert_output_to_messages
+
+ db_messages = [
+ {"id": "m0", "role": "user", "content": "message m0"},
+ {
+ "id": "m1",
+ "role": "assistant",
+ "content": "folded",
+ "output": [{"type": "message", "content": []}],
+ },
+ ]
+ body_messages = [
+ {"role": "user", "content": "message m0"},
+ {"role": "assistant", "content": "converted output"},
+ ]
+
+ self.assertIsNone(
+ self.filter._body_to_db_coverage_map_for_ref_fallback(
+ body_messages,
+ db_messages,
+ )
+ )
+
+ def test_inlet_maps_folded_db_snapshot_to_unfolded_tool_body_tail(self):
+ self.filter.valves.keep_last = 0
+ folded_tool_message = {
+ "id": "m1",
+ "role": "assistant",
+ "content": 'folded result ',
+ "output": [
+ {
+ "role": "assistant",
+ "content": "",
+ "tool_calls": [
+ {
+ "id": "call_1",
+ "type": "function",
+ "function": {"name": "search", "arguments": "{}"},
+ }
+ ],
+ },
+ {
+ "role": "tool",
+ "tool_call_id": "call_1",
+ "content": "large tool result",
+ },
+ {"role": "assistant", "content": "tool follow-up"},
+ ],
+ }
+ db_messages = [
+ {"id": "m0", "role": "user", "content": "message m0"},
+ folded_tool_message,
+ {"id": "m2", "role": "user", "content": "message m2"},
+ {"id": "m3", "role": "assistant", "content": "message m3"},
+ ]
+ body_messages = [
+ {"role": "user", "content": "message m0"},
+ deepcopy(folded_tool_message["output"][0]),
+ {
+ "role": "tool",
+ "tool_call_id": "call_1",
+ "content": "[content collapsed]",
+ "metadata": {
+ "is_trimmed": True,
+ "trimmed_by": "async_context_compression",
+ },
+ },
+ deepcopy(folded_tool_message["output"][2]),
+ {"role": "user", "content": "message m2"},
+ {"role": "assistant", "content": "message m3"},
+ ]
+ snapshots = [
+ _snapshot(
+ "folded tool summary",
+ self.filter._message_refs_for_prefix(db_messages, 2),
+ )
+ ]
+
+ async def fake_load_snapshots(chat_id):
+ return snapshots
+
+ async def fake_load_live_refs(chat_id):
+ return _live_refs_by_id(self.filter, db_messages)
+
+ async def fake_load_full_chat_messages(chat_id):
+ return db_messages
+
+ async def noop(*args, **kwargs):
+ return None
+
+ self.filter._load_summary_snapshots = fake_load_snapshots
+ self.filter._load_chat_history_live_refs = fake_load_live_refs
+ self.filter._load_full_chat_messages = fake_load_full_chat_messages
+ self.filter._log = noop
+ self.filter._emit_debug_log = noop
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 0
+ }
+
+ result = asyncio.run(
+ self.filter.inlet(
+ {
+ "chat_id": "chat-1",
+ "model": "test-model",
+ "messages": body_messages,
+ }
+ )
+ )
+ final_messages = result["messages"]
+
+ self.assertEqual(len(final_messages), 3)
+ self.assertTrue(self.filter._is_summary_message(final_messages[0]))
+ self.assertIn("folded tool summary", final_messages[0]["content"])
+ self.assertEqual(
+ final_messages[0]["metadata"]["covered_until"],
+ 4,
+ )
+ self.assertEqual(
+ [
+ ref["id"]
+ for ref in final_messages[0]["metadata"]["covered_message_refs"]
+ ],
+ ["m0", "m1"],
+ )
+ self.assertEqual(
+ [message["content"] for message in final_messages[1:]],
+ ["message m2", "message m3"],
+ )
+
+ def test_unfolded_db_message_allows_trimmed_assistant_only_with_metadata(self):
+ unfolded_db_message = {
+ "role": "assistant",
+ "content": "full assistant answer with embedded tool details",
+ }
+ trimmed_body_message = {
+ "role": "assistant",
+ "content": "collapsed assistant answer",
+ "metadata": {"tool_outputs_trimmed": True},
+ }
+ unmarked_body_message = {
+ "role": "assistant",
+ "content": "collapsed assistant answer",
+ }
+
+ self.assertTrue(
+ self.filter._body_message_matches_unfolded_db_message(
+ trimmed_body_message,
+ unfolded_db_message,
+ )
+ )
+ self.assertFalse(
+ self.filter._body_message_matches_unfolded_db_message(
+ unmarked_body_message,
+ unfolded_db_message,
+ )
+ )
+
+ def test_inlet_maps_reasoning_output_body_to_folded_db_snapshot(self):
+ self.filter.valves.keep_last = 0
+ misc_module = _ensure_module("open_webui.utils.misc")
+
+ def convert_output_to_messages(output, raw=False):
+ messages = []
+ for item in output:
+ if not isinstance(item, dict) or item.get("type") != "message":
+ continue
+ text = "".join(
+ part.get("text", "")
+ for part in item.get("content", [])
+ if isinstance(part, dict) and part.get("type") == "output_text"
+ )
+ if text:
+ messages.append({"role": "assistant", "content": text})
+ return messages
+
+ misc_module.convert_output_to_messages = convert_output_to_messages
+
+ folded_reasoning_message = {
+ "id": "m1",
+ "role": "assistant",
+ "content": (
+ 'hidden chain \n'
+ "visible answer"
+ ),
+ "output": [
+ {
+ "type": "reasoning",
+ "summary": [{"type": "output_text", "text": "hidden chain"}],
+ },
+ {
+ "type": "message",
+ "content": [
+ {"type": "output_text", "text": "visible answer"}
+ ],
+ },
+ ],
+ }
+ db_messages = [
+ {"id": "m0", "role": "user", "content": "message m0"},
+ folded_reasoning_message,
+ {"id": "m2", "role": "user", "content": "message m2"},
+ {"id": "m3", "role": "assistant", "content": "message m3"},
+ ]
+ body_messages = [
+ {"role": "user", "content": "message m0"},
+ {"role": "assistant", "content": "visible answer"},
+ {"role": "user", "content": "message m2"},
+ {"role": "assistant", "content": "message m3"},
+ ]
+ snapshots = [
+ _snapshot(
+ "folded reasoning summary",
+ self.filter._message_refs_for_prefix(db_messages, 2),
+ )
+ ]
+
+ async def fake_load_snapshots(chat_id):
+ return snapshots
+
+ async def fake_load_live_refs(chat_id):
+ return _live_refs_by_id(self.filter, db_messages)
+
+ async def fake_load_full_chat_messages(chat_id):
+ return db_messages
+
+ async def noop(*args, **kwargs):
+ return None
+
+ self.filter._load_summary_snapshots = fake_load_snapshots
+ self.filter._load_chat_history_live_refs = fake_load_live_refs
+ self.filter._load_full_chat_messages = fake_load_full_chat_messages
+ self.filter._log = noop
+ self.filter._emit_debug_log = noop
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 0
+ }
+
+ result = asyncio.run(
+ self.filter.inlet(
+ {
+ "chat_id": "chat-1",
+ "model": "test-model",
+ "messages": body_messages,
+ }
+ )
+ )
+ final_messages = result["messages"]
+
+ self.assertEqual(len(final_messages), 3)
+ self.assertTrue(self.filter._is_summary_message(final_messages[0]))
+ self.assertIn("folded reasoning summary", final_messages[0]["content"])
+ self.assertEqual(
+ final_messages[0]["metadata"]["covered_until"],
+ 2,
+ )
+ self.assertEqual(
+ [
+ ref["id"]
+ for ref in final_messages[0]["metadata"]["covered_message_refs"]
+ ],
+ ["m0", "m1"],
+ )
+ self.assertEqual(
+ [message["content"] for message in final_messages[1:]],
+ ["message m2", "message m3"],
+ )
+
+ def test_outlet_does_not_reinject_live_sibling_snapshot(self):
+ self.filter.valves.keep_last = 0
+ current_messages = _messages_with_ids(["m0", "m1", "new_m2", "new_m3"])
+ old_branch_messages = _messages_with_ids(["m0", "m1", "old_m2", "old_m3"])
+ old_refs = self.filter._message_refs_for_prefix(old_branch_messages, 4)
+ snapshots = [_snapshot("old branch summary", old_refs)]
+ live_messages = current_messages + old_branch_messages[2:]
+ captured = {}
+ scheduled = []
+
+ async def fake_load_snapshot(
+ chat_id,
+ messages,
+ require_full_coverage=False,
+ ):
+ return self.filter._select_applicable_summary_snapshot(
+ snapshots,
+ messages,
+ require_full_coverage=require_full_coverage,
+ live_message_refs_by_id=_live_refs_by_id(self.filter, live_messages),
+ )
+
+ async def fake_user_context(__user__, __event_call__):
+ return {"user_language": "en-US"}
+
+ async def fake_locked_summary_task(
+ lock,
+ chat_id,
+ model,
+ body,
+ user_data,
+ target_compressed_count,
+ lang,
+ __event_emitter__,
+ __event_call__,
+ __request__=None,
+ ):
+ captured["messages"] = body["messages"]
+
+ async def noop(*args, **kwargs):
+ return None
+
+ def fake_create_task(coro):
+ scheduled.append(coro)
+ return None
+
+ self.filter._load_applicable_summary_snapshot = fake_load_snapshot
+ self.filter._get_user_context = fake_user_context
+ self.filter._get_chat_context = lambda body, metadata=None: {
+ "chat_id": "chat-1",
+ "message_id": "msg-1",
+ }
+ self.filter._should_skip_compression = lambda body, model: False
+ self.filter._locked_summary_task = fake_locked_summary_task
+ self.filter._log = noop
+
+ original_create_task = asyncio.create_task
+ asyncio.create_task = fake_create_task
+ try:
+ asyncio.run(
+ self.filter.outlet(
+ {"model": "test-model", "messages": current_messages},
+ __event_call__=None,
+ )
+ )
+ finally:
+ asyncio.create_task = original_create_task
+
+ self.assertEqual(len(scheduled), 1)
+ asyncio.run(scheduled[0])
+
+ self.assertFalse(
+ any(self.filter._is_summary_message(message) for message in captured["messages"])
+ )
+ self.assertEqual(
+ [message["id"] for message in captured["messages"]],
+ ["m0", "m1", "new_m2", "new_m3"],
+ )
+
+ def test_outlet_reinjects_matching_branch_snapshot_with_metadata(self):
+ self.filter.valves.keep_last = 0
+ current_messages = _messages_with_ids(["m0", "m1", "m2", "m3", "m4"])
+ prefix_refs = self.filter._message_refs_for_prefix(current_messages, 3)
+ snapshots = [_snapshot("shared prefix summary", prefix_refs)]
+ captured = {}
+ scheduled = []
+
+ async def fake_load_snapshot(
+ chat_id,
+ messages,
+ require_full_coverage=False,
+ ):
+ return self.filter._select_applicable_summary_snapshot(
+ snapshots,
+ messages,
+ require_full_coverage=require_full_coverage,
+ live_message_refs_by_id=_live_refs_by_id(self.filter, current_messages),
+ )
+
+ async def fake_user_context(__user__, __event_call__):
+ return {"user_language": "en-US"}
+
+ async def fake_locked_summary_task(
+ lock,
+ chat_id,
+ model,
+ body,
+ user_data,
+ target_compressed_count,
+ lang,
+ __event_emitter__,
+ __event_call__,
+ __request__=None,
+ ):
+ captured["messages"] = body["messages"]
+
+ async def noop(*args, **kwargs):
+ return None
+
+ def fake_create_task(coro):
+ scheduled.append(coro)
+ return None
+
+ self.filter._load_applicable_summary_snapshot = fake_load_snapshot
+ self.filter._get_user_context = fake_user_context
+ self.filter._get_chat_context = lambda body, metadata=None: {
+ "chat_id": "chat-1",
+ "message_id": "msg-1",
+ }
+ self.filter._should_skip_compression = lambda body, model: False
+ self.filter._locked_summary_task = fake_locked_summary_task
+ self.filter._log = noop
+
+ original_create_task = asyncio.create_task
+ asyncio.create_task = fake_create_task
+ try:
+ asyncio.run(
+ self.filter.outlet(
+ {"model": "test-model", "messages": current_messages},
+ __event_call__=None,
+ )
+ )
+ finally:
+ asyncio.create_task = original_create_task
+
+ self.assertEqual(len(scheduled), 1)
+ asyncio.run(scheduled[0])
+ final_messages = captured["messages"]
+
+ self.assertEqual(len(final_messages), 6)
+ self.assertTrue(self.filter._is_summary_message(final_messages[3]))
+ self.assertEqual(
+ [
+ ref["id"]
+ for ref in final_messages[3]["metadata"]["covered_message_refs"]
+ ],
+ ["m0", "m1", "m2"],
+ )
+ self.assertEqual(final_messages[4]["id"], "m3")
+ self.assertEqual(final_messages[5]["id"], "m4")
+
+ def test_snapshot_selection_rejects_same_content_different_ids(self):
+ self.filter.valves.keep_last = 0
+ current_messages = [
+ {"id": "new-1", "role": "user", "content": "same"},
+ {"id": "new-2", "role": "assistant", "content": "same"},
+ ]
+ old_messages = [
+ {"id": "old-1", "role": "user", "content": "same"},
+ {"id": "old-2", "role": "assistant", "content": "same"},
+ ]
+ old_refs = self.filter._message_refs_for_prefix(old_messages, 2)
+
+ selected = self.filter._select_applicable_summary_snapshot(
+ [_snapshot("old same content", old_refs)],
+ current_messages,
+ live_message_refs_by_id=_live_refs_by_id(
+ self.filter, current_messages + old_messages
+ ),
+ )
+
+ self.assertIsNone(selected)
+
+ def test_snapshot_selection_rejects_same_id_changed_payload(self):
+ self.filter.valves.keep_last = 0
+ original_messages = [
+ {"id": "m1", "role": "user", "content": "original question"},
+ {"id": "m2", "role": "assistant", "content": "original answer"},
+ ]
+ edited_messages = [
+ {"id": "m1", "role": "user", "content": "edited question"},
+ {"id": "m2", "role": "assistant", "content": "original answer"},
+ ]
+ original_refs = self.filter._message_refs_for_prefix(original_messages, 2)
+
+ selected = self.filter._select_applicable_summary_snapshot(
+ [_snapshot("old edited content", original_refs)],
+ edited_messages,
+ live_message_refs_by_id=_live_refs_by_id(self.filter, edited_messages),
)
self.assertIsNone(selected)
@@ -2294,6 +2961,84 @@ def get_chat_by_id(chat_id):
self.assertEqual([message["id"] for message in messages], ["m1", "m2", "m3"])
self.assertEqual(messages[2]["role"], "tool")
+ def test_load_full_chat_messages_filters_failed_assistant_from_history_branch(self):
+ class FakeChats:
+ @staticmethod
+ def get_chat_by_id(chat_id):
+ return types.SimpleNamespace(
+ chat={
+ "history": {
+ "currentId": "m4",
+ "messages": {
+ "m1": {
+ "id": "m1",
+ "role": "user",
+ "content": "Question",
+ },
+ "m2": {
+ "id": "m2",
+ "role": "assistant",
+ "content": "",
+ "error": {"message": "provider failed"},
+ "parentId": "m1",
+ },
+ "m3": {
+ "id": "m3",
+ "role": "user",
+ "content": "Retry",
+ "parentId": "m2",
+ },
+ "m4": {
+ "id": "m4",
+ "role": "assistant",
+ "content": "OK",
+ "parentId": "m3",
+ },
+ },
+ }
+ }
+ )
+
+ original_chats = module.Chats
+ module.Chats = FakeChats
+ try:
+ messages = asyncio.run(self.filter._load_full_chat_messages("chat-1"))
+ finally:
+ module.Chats = original_chats
+
+ self.assertEqual([message["id"] for message in messages], ["m1", "m3", "m4"])
+ self.assertFalse(any("error" in message for message in messages))
+
+ def test_load_full_chat_messages_filters_failed_assistant_from_direct_messages(self):
+ class FakeChats:
+ @staticmethod
+ def get_chat_by_id(chat_id):
+ return types.SimpleNamespace(
+ chat={
+ "messages": [
+ {"id": "m1", "role": "user", "content": "Question"},
+ {
+ "id": "m2",
+ "role": "assistant",
+ "content": "",
+ "error": {"message": "provider failed"},
+ },
+ {"id": "m3", "role": "user", "content": "Retry"},
+ {"id": "m4", "role": "assistant", "content": "OK"},
+ ]
+ }
+ )
+
+ original_chats = module.Chats
+ module.Chats = FakeChats
+ try:
+ messages = asyncio.run(self.filter._load_full_chat_messages("chat-1"))
+ finally:
+ module.Chats = original_chats
+
+ self.assertEqual([message["id"] for message in messages], ["m1", "m3", "m4"])
+ self.assertFalse(any("error" in message for message in messages))
+
def test_load_authorized_chat_messages_uses_owner_helper(self):
class FakeChats:
@staticmethod
@@ -3070,6 +3815,7 @@ async def mock_save_summary(
captured["covered_message_refs"] = covered_message_refs
captured["source_current_id"] = source_current_id
captured["protected_head_count"] = protected_head_count
+ return True
async def noop_log(*args, **kwargs):
return None
@@ -3692,16 +4438,108 @@ async def fake_event_call(payload):
else:
module.Users.get_user_by_id = original_get_user
- self.assertIn(
- "Upstream provider error: context too long", str(exc_info.exception)
- )
- self.assertNotIn(
- "LLM response format incorrect or empty", str(exc_info.exception)
- )
- self.assertTrue(frontend_calls)
- self.assertEqual(frontend_calls[0]["type"], "execute")
- self.assertIn("console.error", frontend_calls[0]["data"]["code"])
- self.assertIn("context too long", frontend_calls[0]["data"]["code"])
+ self.assertIn(
+ "Upstream provider error: context too long", str(exc_info.exception)
+ )
+ self.assertNotIn(
+ "LLM response format incorrect or empty", str(exc_info.exception)
+ )
+ self.assertTrue(frontend_calls)
+ self.assertEqual(frontend_calls[0]["type"], "execute")
+ self.assertIn("console.error", frontend_calls[0]["data"]["code"])
+ self.assertIn("context too long", frontend_calls[0]["data"]["code"])
+
+ def test_call_summary_llm_times_out_provider_request(self):
+ self.filter.valves.summary_model = "fake-summary-model"
+ self.filter.valves.max_summary_tokens = 1024
+ self.filter.valves.show_debug_log = False
+ self.filter.valves.summary_fail_mode = "raise"
+ self.filter.valves.summary_llm_timeout_seconds = 0.01
+
+ async def fake_generate_chat_completion(request, payload, user):
+ await asyncio.sleep(10)
+ return {"choices": [{"message": {"content": "too late"}}]}
+
+ async def noop_log(*args, **kwargs):
+ return None
+
+ original_generate = module.generate_chat_completion
+ original_get_user = getattr(module.Users, "get_user_by_id", None)
+
+ module.generate_chat_completion = fake_generate_chat_completion
+ module.Users.get_user_by_id = staticmethod(
+ lambda user_id: types.SimpleNamespace(email="user@example.com")
+ )
+ self.filter._log = noop_log
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 8192
+ }
+ self.filter._build_summary_prompt = (
+ lambda conversation_text, previous_summary=None: conversation_text
+ )
+
+ try:
+ with self.assertRaises(Exception) as exc_info:
+ asyncio.run(
+ self.filter._call_summary_llm(
+ "conversation",
+ {"model": "fake-summary-model"},
+ {"id": "user-1"},
+ )
+ )
+ finally:
+ module.generate_chat_completion = original_generate
+ if original_get_user is None:
+ delattr(module.Users, "get_user_by_id")
+ else:
+ module.Users.get_user_by_id = original_get_user
+
+ self.assertIn("timed out after 0.01 seconds", str(exc_info.exception))
+
+ def test_call_summary_llm_timeout_zero_allows_slow_provider_success(self):
+ self.filter.valves.summary_model = "fake-summary-model"
+ self.filter.valves.max_summary_tokens = 1024
+ self.filter.valves.show_debug_log = False
+ self.filter.valves.summary_llm_timeout_seconds = 0
+
+ async def fake_generate_chat_completion(request, payload, user):
+ await asyncio.sleep(0.01)
+ return {"choices": [{"message": {"content": "eventual summary"}}]}
+
+ async def noop_log(*args, **kwargs):
+ return None
+
+ original_generate = module.generate_chat_completion
+ original_get_user = getattr(module.Users, "get_user_by_id", None)
+
+ module.generate_chat_completion = fake_generate_chat_completion
+ module.Users.get_user_by_id = staticmethod(
+ lambda user_id: types.SimpleNamespace(email="user@example.com")
+ )
+ self.filter._log = noop_log
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 8192
+ }
+ self.filter._build_summary_prompt = (
+ lambda conversation_text, previous_summary=None: conversation_text
+ )
+
+ try:
+ summary = asyncio.run(
+ self.filter._call_summary_llm(
+ "conversation",
+ {"model": "fake-summary-model"},
+ {"id": "user-1"},
+ )
+ )
+ finally:
+ module.generate_chat_completion = original_generate
+ if original_get_user is None:
+ delattr(module.Users, "get_user_by_id")
+ else:
+ module.Users.get_user_by_id = original_get_user
+
+ self.assertEqual(summary, "eventual summary")
def test_extract_summary_text_supports_alternate_response_shapes(self):
self.assertEqual(
@@ -3811,99 +4649,396 @@ async def fake_generate_chat_completion(request, payload, user):
async def noop_log(*args, **kwargs):
return None
- original_generate = module.generate_chat_completion
- original_get_user = getattr(module.Users, "get_user_by_id", None)
-
- module.generate_chat_completion = fake_generate_chat_completion
- module.Users.get_user_by_id = staticmethod(
- lambda user_id: types.SimpleNamespace(email="user@example.com")
- )
+ original_generate = module.generate_chat_completion
+ original_get_user = getattr(module.Users, "get_user_by_id", None)
+
+ module.generate_chat_completion = fake_generate_chat_completion
+ module.Users.get_user_by_id = staticmethod(
+ lambda user_id: types.SimpleNamespace(email="user@example.com")
+ )
+ self.filter._log = noop_log
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 8192
+ }
+ self.filter._build_summary_prompt = (
+ lambda conversation_text, previous_summary=None: conversation_text
+ )
+
+ try:
+ summary = asyncio.run(
+ self.filter._call_summary_llm(
+ "conversation",
+ {"model": "fake-summary-model"},
+ {"id": "user-1"},
+ )
+ )
+ finally:
+ module.generate_chat_completion = original_generate
+ if original_get_user is None:
+ delattr(module.Users, "get_user_by_id")
+ else:
+ module.Users.get_user_by_id = original_get_user
+
+ self.assertEqual(
+ summary,
+ "responses api",
+ )
+
+ def test_call_summary_llm_rejects_empty_message_content(self):
+ self.filter.valves.summary_model = "fake-summary-model"
+ self.filter.valves.max_summary_tokens = 1024
+ self.filter.valves.show_debug_log = False
+ self.filter.valves.summary_fail_mode = "raise"
+
+ async def fake_generate_chat_completion(request, payload, user):
+ return {
+ "choices": [
+ {
+ "message": {
+ "role": "assistant",
+ "content": "",
+ },
+ "finish_reason": "stop",
+ }
+ ]
+ }
+
+ async def noop_log(*args, **kwargs):
+ return None
+
+ original_generate = module.generate_chat_completion
+ original_get_user = getattr(module.Users, "get_user_by_id", None)
+
+ module.generate_chat_completion = fake_generate_chat_completion
+ module.Users.get_user_by_id = staticmethod(
+ lambda user_id: types.SimpleNamespace(email="user@example.com")
+ )
+ self.filter._log = noop_log
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 8192
+ }
+ self.filter._build_summary_prompt = (
+ lambda conversation_text, previous_summary=None: conversation_text
+ )
+
+ try:
+ with self.assertRaises(Exception) as exc_info:
+ asyncio.run(
+ self.filter._call_summary_llm(
+ "conversation",
+ {"model": "fake-summary-model"},
+ {"id": "user-1"},
+ )
+ )
+ finally:
+ module.generate_chat_completion = original_generate
+ if original_get_user is None:
+ delattr(module.Users, "get_user_by_id")
+ else:
+ module.Users.get_user_by_id = original_get_user
+
+ self.assertIn(
+ "LLM response did not contain summary text", str(exc_info.exception)
+ )
+
+ def test_generate_summary_async_status_guides_user_to_browser_console(self):
+ self.filter.valves.keep_first = 1
+ self.filter.valves.keep_last = 1
+ self.filter.valves.summary_model = "fake-summary-model"
+ self.filter.valves.summary_model_max_context = 1200
+ self.filter.valves.max_summary_tokens = 500
+ self.filter.valves.show_debug_log = False
+
+ events = []
+ frontend_calls = []
+
+ async def fake_summary_llm(*args, **kwargs):
+ raise Exception("boom details")
+
+ async def fake_emitter(event):
+ events.append(event)
+
+ async def fake_event_call(payload):
+ frontend_calls.append(payload)
+ return True
+
+ async def noop_log(*args, **kwargs):
+ return None
+
+ self.filter._log = noop_log
+ self.filter._call_summary_llm = fake_summary_llm
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 1200
+ }
+ self.filter._format_messages_for_summary = lambda messages: "\n".join(
+ msg["content"] for msg in messages
+ )
+ self.filter._build_summary_prompt = (
+ lambda conversation_text, previous_summary=None: conversation_text
+ )
+ self.filter._count_tokens = lambda text: len(text)
+
+ messages = [
+ {"role": "system", "content": "System prompt"},
+ {"role": "user", "content": "Q" * 40},
+ {"role": "assistant", "content": "A" * 40},
+ {"role": "user", "content": "Question 2"},
+ ]
+
+ asyncio.run(
+ self.filter._generate_summary_async(
+ messages=messages,
+ chat_id="chat-1",
+ body={"model": "fake-summary-model"},
+ user_data={"id": "user-1"},
+ target_compressed_count=3,
+ lang="en-US",
+ __event_emitter__=fake_emitter,
+ __event_call__=fake_event_call,
+ )
+ )
+
+ self.assertTrue(frontend_calls)
+ self.assertIn("console.error", frontend_calls[0]["data"]["code"])
+ self.assertIn("boom details", frontend_calls[0]["data"]["code"])
+ status_descriptions = [
+ event["data"]["description"]
+ for event in events
+ if event.get("type") == "status"
+ ]
+ self.assertTrue(
+ any("Check browser console (F12) for details" in text for text in status_descriptions)
+ )
+
+ def test_generate_summary_async_empty_summary_settles_generating_status(self):
+ self.filter.valves.keep_first = 1
+ self.filter.valves.keep_last = 1
+ self.filter.valves.summary_model = "fake-summary-model"
+ self.filter.valves.summary_model_max_context = 1200
+ self.filter.valves.max_summary_tokens = 500
+ self.filter.valves.show_debug_log = False
+
+ events = []
+ save_called = False
+
+ async def empty_summary_llm(*args, **kwargs):
+ return ""
+
+ async def fake_save_summary(*args, **kwargs):
+ nonlocal save_called
+ save_called = True
+ return True
+
+ async def fake_emitter(event):
+ events.append(event)
+
+ async def no_snapshot(*args, **kwargs):
+ return None
+
+ async def noop_log(*args, **kwargs):
+ return None
+
+ self.filter._log = noop_log
+ self.filter._call_summary_llm = empty_summary_llm
+ self.filter._save_summary = fake_save_summary
+ self.filter._load_applicable_summary_snapshot = no_snapshot
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 1200
+ }
+ self.filter._format_messages_for_summary = lambda messages: "\n".join(
+ msg["content"] for msg in messages
+ )
+ self.filter._build_summary_prompt = (
+ lambda conversation_text, previous_summary=None: conversation_text
+ )
+ self.filter._count_tokens = lambda text: len(text)
+
+ messages = [
+ {"id": "m0", "role": "system", "content": "System prompt"},
+ {"id": "m1", "role": "user", "content": "Q" * 40},
+ {"id": "m2", "role": "assistant", "content": "A" * 40},
+ {"id": "m3", "role": "user", "content": "Question 2"},
+ ]
+
+ asyncio.run(
+ self.filter._generate_summary_async(
+ messages=messages,
+ chat_id="chat-1",
+ body={"model": "fake-summary-model"},
+ user_data={"id": "user-1"},
+ target_compressed_count=3,
+ lang="en-US",
+ __event_emitter__=fake_emitter,
+ __event_call__=None,
+ )
+ )
+
+ statuses = [
+ event["data"] for event in events if event.get("type") == "status"
+ ]
+ self.assertGreaterEqual(len(statuses), 2)
+ self.assertEqual(
+ statuses[-2]["description"], "Generating context summary in background..."
+ )
+ self.assertFalse(statuses[-2]["done"])
+ self.assertTrue(statuses[-1]["done"])
+ self.assertIn("Summary Error", statuses[-1]["description"])
+ self.assertIn("empty", statuses[-1]["description"])
+ self.assertFalse(save_called)
+
+ def test_generate_summary_async_save_failure_settles_generating_status(self):
+ self.filter.valves.keep_first = 1
+ self.filter.valves.keep_last = 1
+ self.filter.valves.summary_model = "fake-summary-model"
+ self.filter.valves.summary_model_max_context = 1200
+ self.filter.valves.max_summary_tokens = 500
+ self.filter.valves.show_debug_log = False
+
+ events = []
+ save_called = False
+
+ async def fake_summary_llm(*args, **kwargs):
+ return "new summary"
+
+ async def fail_save_summary(*args, **kwargs):
+ nonlocal save_called
+ save_called = True
+ return False
+
+ async def fake_emitter(event):
+ events.append(event)
+
+ async def no_snapshot(*args, **kwargs):
+ return None
+
+ async def noop_log(*args, **kwargs):
+ return None
+
self.filter._log = noop_log
+ self.filter._call_summary_llm = fake_summary_llm
+ self.filter._save_summary = fail_save_summary
+ self.filter._load_applicable_summary_snapshot = no_snapshot
self.filter._get_model_thresholds = lambda model_id: {
- "max_context_tokens": 8192
+ "max_context_tokens": 1200
}
+ self.filter._format_messages_for_summary = lambda messages: "\n".join(
+ msg["content"] for msg in messages
+ )
self.filter._build_summary_prompt = (
lambda conversation_text, previous_summary=None: conversation_text
)
+ self.filter._count_tokens = lambda text: len(text)
- try:
- summary = asyncio.run(
- self.filter._call_summary_llm(
- "conversation",
- {"model": "fake-summary-model"},
- {"id": "user-1"},
- )
+ messages = [
+ {"id": "m0", "role": "system", "content": "System prompt"},
+ {"id": "m1", "role": "user", "content": "Q" * 40},
+ {"id": "m2", "role": "assistant", "content": "A" * 40},
+ {"id": "m3", "role": "user", "content": "Question 2"},
+ ]
+
+ asyncio.run(
+ self.filter._generate_summary_async(
+ messages=messages,
+ chat_id="chat-1",
+ body={"model": "fake-summary-model"},
+ user_data={"id": "user-1"},
+ target_compressed_count=3,
+ lang="en-US",
+ __event_emitter__=fake_emitter,
+ __event_call__=None,
)
- finally:
- module.generate_chat_completion = original_generate
- if original_get_user is None:
- delattr(module.Users, "get_user_by_id")
- else:
- module.Users.get_user_by_id = original_get_user
+ )
+ statuses = [
+ event["data"] for event in events if event.get("type") == "status"
+ ]
+ self.assertGreaterEqual(len(statuses), 2)
self.assertEqual(
- summary,
- "responses api",
+ statuses[-2]["description"], "Generating context summary in background..."
+ )
+ self.assertFalse(statuses[-2]["done"])
+ self.assertTrue(statuses[-1]["done"])
+ self.assertIn("Summary Error", statuses[-1]["description"])
+ self.assertIn("persisted", statuses[-1]["description"])
+ self.assertTrue(save_called)
+ self.assertFalse(
+ any("Loaded historical summary" in status["description"] for status in statuses)
)
- def test_call_summary_llm_rejects_empty_message_content(self):
+ def test_generate_summary_async_terminal_status_emitter_failure_is_best_effort(self):
+ self.filter.valves.keep_first = 1
+ self.filter.valves.keep_last = 1
self.filter.valves.summary_model = "fake-summary-model"
- self.filter.valves.max_summary_tokens = 1024
+ self.filter.valves.summary_model_max_context = 1200
+ self.filter.valves.max_summary_tokens = 500
self.filter.valves.show_debug_log = False
- self.filter.valves.summary_fail_mode = "raise"
- async def fake_generate_chat_completion(request, payload, user):
- return {
- "choices": [
- {
- "message": {
- "role": "assistant",
- "content": "",
- },
- "finish_reason": "stop",
- }
- ]
- }
+ events = []
+ save_called = False
- async def noop_log(*args, **kwargs):
+ async def empty_summary_llm(*args, **kwargs):
+ return ""
+
+ async def fake_save_summary(*args, **kwargs):
+ nonlocal save_called
+ save_called = True
+ return True
+
+ async def flaky_emitter(event):
+ if events:
+ raise RuntimeError("frontend disconnected")
+ events.append(event)
+
+ async def no_snapshot(*args, **kwargs):
return None
- original_generate = module.generate_chat_completion
- original_get_user = getattr(module.Users, "get_user_by_id", None)
+ async def noop_log(*args, **kwargs):
+ return None
- module.generate_chat_completion = fake_generate_chat_completion
- module.Users.get_user_by_id = staticmethod(
- lambda user_id: types.SimpleNamespace(email="user@example.com")
- )
self.filter._log = noop_log
+ self.filter._call_summary_llm = empty_summary_llm
+ self.filter._save_summary = fake_save_summary
+ self.filter._load_applicable_summary_snapshot = no_snapshot
self.filter._get_model_thresholds = lambda model_id: {
- "max_context_tokens": 8192
+ "max_context_tokens": 1200
}
+ self.filter._format_messages_for_summary = lambda messages: "\n".join(
+ msg["content"] for msg in messages
+ )
self.filter._build_summary_prompt = (
lambda conversation_text, previous_summary=None: conversation_text
)
+ self.filter._count_tokens = lambda text: len(text)
- try:
- with self.assertRaises(Exception) as exc_info:
- asyncio.run(
- self.filter._call_summary_llm(
- "conversation",
- {"model": "fake-summary-model"},
- {"id": "user-1"},
- )
- )
- finally:
- module.generate_chat_completion = original_generate
- if original_get_user is None:
- delattr(module.Users, "get_user_by_id")
- else:
- module.Users.get_user_by_id = original_get_user
+ messages = [
+ {"id": "m0", "role": "system", "content": "System prompt"},
+ {"id": "m1", "role": "user", "content": "Q" * 40},
+ {"id": "m2", "role": "assistant", "content": "A" * 40},
+ {"id": "m3", "role": "user", "content": "Question 2"},
+ ]
- self.assertIn(
- "LLM response did not contain summary text", str(exc_info.exception)
+ asyncio.run(
+ self.filter._generate_summary_async(
+ messages=messages,
+ chat_id="chat-1",
+ body={"model": "fake-summary-model"},
+ user_data={"id": "user-1"},
+ target_compressed_count=3,
+ lang="en-US",
+ __event_emitter__=flaky_emitter,
+ __event_call__=None,
+ )
)
- def test_generate_summary_async_status_guides_user_to_browser_console(self):
+ self.assertEqual(len(events), 1)
+ self.assertEqual(
+ events[0]["data"]["description"],
+ "Generating context summary in background...",
+ )
+ self.assertFalse(events[0]["data"]["done"])
+ self.assertFalse(save_called)
+
+ def test_generate_summary_async_error_status_emitter_failure_is_best_effort(self):
self.filter.valves.keep_first = 1
self.filter.valves.keep_last = 1
self.filter.valves.summary_model = "fake-summary-model"
@@ -3912,23 +5047,24 @@ def test_generate_summary_async_status_guides_user_to_browser_console(self):
self.filter.valves.show_debug_log = False
events = []
- frontend_calls = []
- async def fake_summary_llm(*args, **kwargs):
- raise Exception("boom details")
+ async def fail_summary_llm(*args, **kwargs):
+ raise Exception("summary backend failed")
- async def fake_emitter(event):
+ async def flaky_emitter(event):
+ if events:
+ raise RuntimeError("frontend disconnected")
events.append(event)
- async def fake_event_call(payload):
- frontend_calls.append(payload)
- return True
+ async def no_snapshot(*args, **kwargs):
+ return None
async def noop_log(*args, **kwargs):
return None
self.filter._log = noop_log
- self.filter._call_summary_llm = fake_summary_llm
+ self.filter._call_summary_llm = fail_summary_llm
+ self.filter._load_applicable_summary_snapshot = no_snapshot
self.filter._get_model_thresholds = lambda model_id: {
"max_context_tokens": 1200
}
@@ -3941,10 +5077,10 @@ async def noop_log(*args, **kwargs):
self.filter._count_tokens = lambda text: len(text)
messages = [
- {"role": "system", "content": "System prompt"},
- {"role": "user", "content": "Q" * 40},
- {"role": "assistant", "content": "A" * 40},
- {"role": "user", "content": "Question 2"},
+ {"id": "m0", "role": "system", "content": "System prompt"},
+ {"id": "m1", "role": "user", "content": "Q" * 40},
+ {"id": "m2", "role": "assistant", "content": "A" * 40},
+ {"id": "m3", "role": "user", "content": "Question 2"},
]
asyncio.run(
@@ -3955,22 +5091,108 @@ async def noop_log(*args, **kwargs):
user_data={"id": "user-1"},
target_compressed_count=3,
lang="en-US",
- __event_emitter__=fake_emitter,
- __event_call__=fake_event_call,
+ __event_emitter__=flaky_emitter,
+ __event_call__=None,
)
)
- self.assertTrue(frontend_calls)
- self.assertIn("console.error", frontend_calls[0]["data"]["code"])
- self.assertIn("boom details", frontend_calls[0]["data"]["code"])
- status_descriptions = [
- event["data"]["description"]
- for event in events
- if event.get("type") == "status"
+ self.assertEqual(len(events), 1)
+ self.assertEqual(
+ events[0]["data"]["description"],
+ "Generating context summary in background...",
+ )
+ self.assertFalse(events[0]["data"]["done"])
+
+ def test_generate_summary_async_timeout_silent_settles_generating_status(self):
+ self.filter.valves.keep_first = 1
+ self.filter.valves.keep_last = 1
+ self.filter.valves.summary_model = "fake-summary-model"
+ self.filter.valves.summary_model_max_context = 1200
+ self.filter.valves.max_summary_tokens = 500
+ self.filter.valves.show_debug_log = False
+ self.filter.valves.summary_llm_timeout_seconds = 0.01
+
+ events = []
+ save_called = False
+
+ async def slow_generate_chat_completion(request, payload, user):
+ await asyncio.sleep(10)
+ return {"choices": [{"message": {"content": "too late"}}]}
+
+ async def fake_save_summary(*args, **kwargs):
+ nonlocal save_called
+ save_called = True
+ return True
+
+ async def fake_emitter(event):
+ events.append(event)
+
+ async def no_snapshot(*args, **kwargs):
+ return None
+
+ async def noop_log(*args, **kwargs):
+ return None
+
+ original_generate = module.generate_chat_completion
+ original_get_user = getattr(module.Users, "get_user_by_id", None)
+
+ module.generate_chat_completion = slow_generate_chat_completion
+ module.Users.get_user_by_id = staticmethod(
+ lambda user_id: types.SimpleNamespace(email="user@example.com")
+ )
+ self.filter._log = noop_log
+ self.filter._save_summary = fake_save_summary
+ self.filter._load_applicable_summary_snapshot = no_snapshot
+ self.filter._get_model_thresholds = lambda model_id: {
+ "max_context_tokens": 1200
+ }
+ self.filter._format_messages_for_summary = lambda messages: "\n".join(
+ msg["content"] for msg in messages
+ )
+ self.filter._build_summary_prompt = (
+ lambda conversation_text, previous_summary=None: conversation_text
+ )
+ self.filter._count_tokens = lambda text: len(text)
+
+ messages = [
+ {"id": "m0", "role": "system", "content": "System prompt"},
+ {"id": "m1", "role": "user", "content": "Q" * 40},
+ {"id": "m2", "role": "assistant", "content": "A" * 40},
+ {"id": "m3", "role": "user", "content": "Question 2"},
]
- self.assertTrue(
- any("Check browser console (F12) for details" in text for text in status_descriptions)
+
+ try:
+ asyncio.run(
+ self.filter._generate_summary_async(
+ messages=messages,
+ chat_id="chat-1",
+ body={"model": "fake-summary-model"},
+ user_data={"id": "user-1"},
+ target_compressed_count=3,
+ lang="en-US",
+ __event_emitter__=fake_emitter,
+ __event_call__=None,
+ )
+ )
+ finally:
+ module.generate_chat_completion = original_generate
+ if original_get_user is None:
+ delattr(module.Users, "get_user_by_id")
+ else:
+ module.Users.get_user_by_id = original_get_user
+
+ statuses = [
+ event["data"] for event in events if event.get("type") == "status"
+ ]
+ self.assertGreaterEqual(len(statuses), 2)
+ self.assertEqual(
+ statuses[-2]["description"], "Generating context summary in background..."
)
+ self.assertFalse(statuses[-2]["done"])
+ self.assertTrue(statuses[-1]["done"])
+ self.assertIn("Summary Error", statuses[-1]["description"])
+ self.assertIn("empty", statuses[-1]["description"])
+ self.assertFalse(save_called)
def test_check_and_generate_summary_async_forces_frontend_and_status_on_pre_summary_error(
self,
@@ -4025,6 +5247,55 @@ def fail_estimate(_messages):
any("Check browser console (F12) for details" in text for text in status_descriptions)
)
+ def test_check_and_generate_summary_async_context_status_emitter_failure_still_generates_summary(
+ self,
+ ):
+ self.filter.valves.show_debug_log = False
+ self.filter.valves.show_token_usage_status = True
+ self.filter.valves.token_usage_status_threshold = 0
+
+ generated = False
+ events = []
+
+ async def fake_generate_summary_async(*args, **kwargs):
+ nonlocal generated
+ generated = True
+
+ async def flaky_emitter(event):
+ events.append(event)
+ if len(events) == 1:
+ raise RuntimeError("frontend disconnected")
+
+ async def noop_log(*args, **kwargs):
+ return None
+
+ self.filter._log = noop_log
+ self.filter._generate_summary_async = fake_generate_summary_async
+ self.filter._estimate_messages_tokens = lambda messages: 100
+ self.filter._calculate_messages_tokens = lambda messages: 100
+ self.filter._get_model_thresholds = lambda model_id: {
+ "compression_threshold_tokens": 100,
+ "max_context_tokens": 1000,
+ }
+
+ asyncio.run(
+ self.filter._check_and_generate_summary_async(
+ chat_id="chat-1",
+ model="fake-model",
+ body={"messages": [{"role": "user", "content": "Hello"}]},
+ user_data={"id": "user-1"},
+ target_compressed_count=1,
+ lang="en-US",
+ __event_emitter__=flaky_emitter,
+ __event_call__=None,
+ )
+ )
+
+ self.assertTrue(generated)
+ self.assertEqual(len(events), 1)
+ self.assertIn("Context Usage", events[0]["data"]["description"])
+ self.assertTrue(events[0]["data"]["done"])
+
def test_external_reference_message_detection_matches_injected_marker(self):
message = {
"role": "assistant",
diff --git a/plugins/filters/async-context-compression/v1.7.1.md b/plugins/filters/async-context-compression/v1.7.1.md
new file mode 100644
index 0000000..fadabfd
--- /dev/null
+++ b/plugins/filters/async-context-compression/v1.7.1.md
@@ -0,0 +1,16 @@
+# Async Context Compression v1.7.1 Release Notes
+
+## Overview
+
+This patch release fixes summary reuse for OpenWebUI request bodies that do not expose stable message ids. The filter can now prove those requests against the persisted DB active branch before falling back to raw history, which prevents very long chats from sending hundreds of thousands of raw tokens when a branch-valid summary already exists.
+
+## Fixes
+
+- **Idless request body matching**: If the model-visible request body lacks message ids, the filter can compare it with the persisted active branch and reuse DB-backed refs only when the visible payloads match.
+- **Terminal assistant placeholder handling**: If OpenWebUI has inserted an in-progress assistant message at the end of the DB branch, the filter may ignore that terminal assistant only when the request body proves it matches the branch ending at the latest user message.
+- **Folded output and failed assistant parity**: DB branch matching now accounts for folded assistant output, model-visible output expansion, and failed assistant messages that middleware filters out.
+- **Safer fallback behavior**: Folded output conversion errors fail closed, and debug logging no longer crashes idless fallback selection.
+
+## Upgrade Notes
+
+No database migration is required beyond the v1.7.0 branch-aware schema. Update or reinstall the filter so OpenWebUI's stored function content includes the v1.7.1 matching and fallback fixes.
diff --git a/plugins/filters/async-context-compression/v1.7.1_CN.md b/plugins/filters/async-context-compression/v1.7.1_CN.md
new file mode 100644
index 0000000..2742b89
--- /dev/null
+++ b/plugins/filters/async-context-compression/v1.7.1_CN.md
@@ -0,0 +1,16 @@
+# 异步上下文压缩 v1.7.1 版本发布说明
+
+## 概述
+
+这个补丁版本修复了 OpenWebUI request body 没有稳定 message id 时的摘要复用问题。插件现在会先用数据库里的 active branch 严格证明当前请求对应同一条可见分支,再决定是否复用已有摘要,避免长对话在已有 branch-valid 摘要时误退回几十万 token 的原始历史。
+
+## 修复内容
+
+- **无 id 请求体匹配**:如果模型可见 request body 缺少 message id,过滤器会和持久化 active branch 对齐比较;只有可见 payload 匹配时,才使用 DB refs 复用摘要。
+- **末尾 assistant 占位消息处理**:如果 OpenWebUI 已经把“正在生成中”的 assistant 消息写到 DB 分支末尾,插件只有在 request body 能证明匹配到最新 user message 处的分支时,才会忽略这个末尾 assistant。
+- **folded output 和失败 assistant 对齐**:DB 分支匹配现在会考虑折叠的 assistant output、模型可见 output 展开,以及 middleware 会过滤掉的失败 assistant 消息。
+- **更安全的 fallback 行为**:folded output 转换异常会 fail closed;debug 日志也不会再让 idless fallback 选择路径崩溃。
+
+## 升级说明
+
+除了 v1.7.0 的 branch-aware schema 之外,本版本不需要新的数据库迁移。请更新或重新安装过滤器,确保 OpenWebUI 中保存的 function 内容包含 v1.7.1 的匹配和 fallback 修复。