diff --git a/CHANGELOG.md b/CHANGELOG.md index 99947cd7..e2f39448 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,3 +1,64 @@ +# v5.5.2 + +### Selfmedia Operator 视频制作与分发能力 + +- **一站式短视频制作**:`video-product` 技能支持文章链接、追爆报告、文字主题、本地文件等多种输入,自动生成脚本 → 逐段生成视频素材(声画同出)→ FFmpeg 合成成片。直连火山引擎 Seedance(doubao-seedance-2.0 系列)与阿里云百炼 Wan2.7-HappyHorse(happyhorse-1.1 系列)端点,按平台自动 fallback +- **视频分发**:新增微信视频号发布(`wechat-channels-publish`,处理 wujie shadow DOM),结合既有的小红书、抖音、Twitter/X、B站、快手等平台,实现短视频制作 → 多平台分发的闭环 +- **两个剪辑辅助技能**: + - `de-mouth`:口播视频去口误,自动识别并删除静音、语气词、卡顿词、重复句、残句,输出干净视频 + 字幕 + 剪映草稿 + - `highlight-clipper`:从本地视频中通过 ASR 转录 + 文本分析自动提取高光片段,剪辑输出多段短视频 +- Selfmedia Operator引入科学的评估方案和自动复盘方案(发布前预测打分 -> 每日数据复盘 -> 根据复盘调整打分量表 -> 不断优化预测准确性)。以上已内置到所有平台的发布流程中,让运营工作不再“凭感觉”。 + +### Smart Search重构 +- 采用路由模式。 + +### 主力模型切换为 GLM-5.2,推荐火山方舟 Coding Plan + +- `config-templates/openclaw.json` 主力模型由 DeepSeek V4 Pro 切换为 **GLM-5.2**(经火山引擎方舟 Coding Plan 接入,`awk/glm-latest`),fallback 为 siliconflow provider +- `install.sh` 交互式收集的 key 由 `DEEPSEEK_API_KEY` 改为 `AWK_API_KEY` +- 大模型推荐主推**火山方舟 Coding Plan**:支持 GLM-5.2、Kimi-K2.7、MiniMax-M3、DeepSeek-V4 系列、Doubao-Seed-2.0 系列等模型,工具不限;通过 wiseflow 邀请链接订阅叠加 9.5 折,首月尝鲜低至 9.4 元。邀请链接 https://volcengine.com/L/dx-wt80li-I/ ,邀请码 `5Y5A6L86` +- siliconflow、aihubmix 推荐不变(siliconflow 仍需申请,作为视觉/替补模型) + +> 想使用 5.5.2 的视频生成能力,需额外开通火山方舟 doubao-seedance-2.0 系列或阿里云百炼 happyhorse-1.1 系列模型,并将对应 key(`AWK_GEN_KEY` 或 `MODELSTUDIO_API_KEY`)配置到 `daemon.env`。 + +### openclaw 上游同步至 v2026.6.10 + +- 从 v2026.6.6 升级到 v2026.6.10 +- **删除 patch 001**(relax exec allowlist shell syntax):上游 exec 审批重构为 risk-based(`command-explainer` + `exec-authorization-plan`),`&&`/`||`/`;` 复合命令已原生逐段匹配 allowlist;`$()`/反引号/重定向上游仍拒但 wiseflow 已改走 `.sh` 脚本。原目标代码 `splitShellPipeline` 已删,无法 re-port +- **删除 patch 004**(chrome port grace retry):上游新增 `ensureManagedChromePortAvailable` + `recoverOwnedStaleManagedChromeCdpListener`,命中 EADDRINUSE 时主动杀掉占用端口的陈旧 Chrome 进程并清 singleton lock 再重探,比 3×500ms 轮询更强 +- 保留 patch 002/003/005/006(验证 apply 通过,上游无等价改动) + +### 上游关键变更摘要(与 wiseflow 相关) + +- **GLM-5.2(6.10)**:暴露 reasoning levels、GLM overload failover、Zai 合成模型回退 manifest baseUrl +- **心跳(6.9)**:修复 5.20 及所有 5.x 上心跳 scheduler 不触发的回归(#88970) +- **sessions_yield over MCP(6.9,#90861)**:修复 MCP 下 sessions_yield 保留 +- **安全(6.9)**:secrets redaction、阻断内部 HTTP session overrides、审计 open-DM tool exposure、plugin write owner check +- **存储(6.9)**:NFS 上禁用 SQLite WAL、reindex temp 清理、setup state 移出 workspace dot-dir +- **web search(6.9)**:Codex Hosted Search、key-free provider 保持 opt-in +- **6.10**:fast talks auto mode、channel switch reset 陈旧 origin 字段、hook registry 组合保留 trusted policies + +# v5.5.1 + +### openclaw 上游同步至 v2026.6.6 + +- 从 v2026.5.28 升级到 v2026.6.6(4253 commits,跨越 6.1→6.2→6.5→6.6 四个稳定版) +- 重新生成 patch 004(chrome-port-grace-retry)和 patch 005(browser-timeout-env-var)以适配上游文件重构 +- patch 001–003 验证通过,无需修改 + +### 上游关键变更摘要 + +- **安全加固**:exec 审批超时默认拒绝(fail-closed)、sandbox binds 收紧、MCP stdio 继承收紧、Codex HTTP 私有目标阻断、loopback tools 权限隔离 +- **OpenRouter 一等公民**:模型设置流程原生支持 OpenRouter OAuth/API-key +- **Parallel Search (Free)**:零配置内置 web search(无需 API key),作为 DuckDuckGo 之前的默认 fallback +- **移动端**:iPad 侧边栏 + iPhone Control Hub,Workboard/Skill Workshop 连接 Gateway +- **Telegram/iMessage**:account-scoped topic 路由、always-on inbound restart、durable echo markers +- **Browser/MCP**:existing-session CDP 支持、WebSocket validation、Streamable HTTP loopback、OAuth/SSE auth 修正 +- **Provider**:Claude Fable 5 adaptive thinking、Gemma 4 reasoning replay、本地模型跳过 guardian review、gpt-5.3-codex 恢复 +- **Cron**:wake 保留 originating session/agent、impossible cron 表达式拒绝创建 +- **启动提速**:cached model metadata、移除 startup catalog wait、lazy slash-command loading +- **QoL**:`openclaw update repair` 恢复路径、compaction timeout 默认降至 180s + # v5.5.0 ### 完全重新设计的部署与渠道绑定流程 diff --git a/CLAUDE.md b/CLAUDE.md index d9a867c6..7f188e6f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -30,11 +30,7 @@ Claude Code 被授权在本仓库中执行任何 git 命令(包括 push、bran - 1、多步骤操作且涉及中间态保存的(下一步操作的某一输入为上一步返回结果),哪怕每一步都只是一条命令,也必须做脚本! - 2、涉及多分支选择,且分支选择依靠明确变量的(如环境变量中是否有某个值,或者按某个入参的值判断分支)应该优先用脚本。 - 3、涉及 python 的,必须制作脚本,最终以 “python /path/to/script.py” 的模式调用。 -- 4、**crew 专属 skill**(`crews/` 或 `addons/officials/crew/` 下的 skill)如果包含脚本,SKILL.md 中对脚本调用的路径必须使用相对路径写法,即 `./skills//scripts/`,**不得**使用 `{baseDir}/scripts/...`。 - -原因:openclaw exec allowlist 以 workspace 为 CWD 做相对路径匹配;`{baseDir}` 是 claude code 专用变量,在 openclaw 中不会展开。全局 skill(`skills/` 目录下)不受此限制,使用 `{baseDir}` 即可。 - -- 5、skill 需要的常量(如各种 ID、KEY 等),搭配脚本时优先使用环境变量,搭配 SKILL.md 时优先使用同级目录下的 json 配置。 +- 4、skill 需要的常量(如各种 ID、KEY 等),搭配脚本时优先使用环境变量,搭配 SKILL.md 时优先使用同级目录下的 json 配置。 本代码仓的 skill 是给 openclaw 使用的,以上原则是为了适配 openclaw 的规则。 diff --git a/README.md b/README.md index 59e84beb..5032a09c 100644 --- a/README.md +++ b/README.md @@ -1,14 +1,14 @@ # Wiseflow -🚀 **v5.5.0 更新** +🚀 **v5.5.2 更新** -- **从安装到出活,全程微信对话完成**:首次部署时自动安装官方微信插件,扫码绑定后,Wiseflow Main Agent 直接在微信上引导你完成业务背景采集、团队组建、渠道配置——不再需要编辑任何配置文件 -- 如果只是想要一个个人助理、或者使用场景比较简单(crew 数量不超过 3 个),可以一直使用微信渠道,无需额外申请飞书开放平台或企业微信(后续如需扩展,Main Agent 也会给出申请和开通的详细指导) -- **Selfmedia Operator 增强**:新增微信公众号文章自动排版并推送至草稿箱、简单短视频制作(t2video)、高光时刻视频剪辑(highlight-clipper),现已支持 15 个国内外主流自媒体平台的发布能力(部分为pro版本提供) -- **Designer 升级**:从"出图工具"重新定位为**系统性视觉设计体系构建者**,内置 `design-system-picker` 技能,预置 15 套品牌设计系统(覆盖 fintech / devtools / productivity / consumer / luxury / enterprise 等全品类),一键匹配风格后从零构建完整网页、APP 界面、品牌视觉体系 -- **IT Engineer 升级**:新增SEO、ICP备案辅助、云服务资源管理能力,现在 Designer + IT Engineer 的技能组合可覆盖官网 / Landing Page 的完整流程——**设计 → 开发 → 部署(云计算)→ 备案(ICP)→ SEO** -- **new crew: business-developer**:business-developer 继承 4.x 全部核心能力(指定信源监控、行业情报采集、社交媒体潜客挖掘),并新增业务介绍 PPT 制作能力 -- **new crew: investor-relationship**(预发布):软件著作权申报(材料准备、页面填写)、商业模式优化、商业思路记录、BP编写、项目申报、搜索并接触投资人 +- **Selfmedia Operator 新增视频制作与分发能力**:一站式短视频制作(`video-product`)支持根据一个主题或一篇文章的链接全流程端到端制作短视频:自动生成脚本 → 逐段生成视频素材(声画同出)→ 合成成片 +- Selfmedia Operator 打通微信视频号、小红书、抖音、Twitter/X 等平台分发,形成"制作 → 分发"闭环 +- 为Selfmedia Operator引入科学的评估方案和自动复盘方案(发布前预测打分 -> 每日数据复盘 -> 根据复盘调整打分量表 -> 不断优化预测准确性)。以上已内置到所有平台的发布流程中,让运营工作不再“凭感觉”。 +- **两个剪辑辅助技能**:`de-mouth`(口播视频去口误,自动识别并删除静音、语气词、卡顿词、重复句、残句)、`highlight-clipper`(ASR + 文本分析自动提取高光片段剪成多段短视频) +- Smart Search技能重构,采用灵活路由机制。 +- **主力模型切换为 GLM-5.2,推荐火山方舟 Coding Plan**:通过 wiseflow 邀请链接订阅叠加 9.5 折,首月低至 9.4 元 +- 适配openclaw 2026-6-10 版本、openclaw-weixin 2.4.6版本 详见 [CHANGELOG.md](CHANGELOG.md) @@ -18,6 +18,8 @@ Wiseflow 是基于 [openclaw](https://github.com/openclaw/openclaw) 的 Multi-Agent 系统,为 **所有被/或即将为 AI 时代冲击、需要独立拓展收入来源的个体** 打造——被裁员/降薪的职场人、副业探索者、自媒体个体户、小生意人、刚毕业的年轻人…… +**最小安装后全程仅需微信即可使用,无需额外配置其他软件**,零安装负担! + > **对于 99% 的人来说,人工智能技术带来的其实是灾难** > > 这不是危言耸听。历史上每一次技术变革——蒸汽机、电力、互联网——无一例外都让会用它的人赚得更多,不会用的人被甩得更远。因为技术本质上是杠杆:有资本、有资源的人能第一时间装备自己,效率翻倍;而普通人连反应的时间都没有,就已经被替代了。AI 时代只会把这个规律放大到极致——99% 的输家,1% 的赢家。这非常不公平、也不合理。然而遗憾的是,这场变革已经无法被停止,那么,我们能做点什么? @@ -40,11 +42,13 @@ wiseflow 目前能为你提供的是: ### 0. 准备 API Key -1. 注册 [DeepSeek 官方 API](https://platform.deepseek.com/) 并充值(前期试水,充个 10 块钱够了),获得 `DEEPSEEK_API_KEY` -2. 注册 [SiliconFlow](https://cloud.siliconflow.cn/i/WNLYbBpi)(🎁 欢迎使用我的邀请链接,你我均会获得 16 元代金券),获得 `SILICONFLOW_API_KEY` +1. 注册 [火山引擎方舟 Coding Plan](https://volcengine.com/L/dx-wt80li-I/)(🎁 欢迎使用 wiseflow 邀请链接 / 邀请码 `5Y5A6L86`,订阅叠加 9.5 折,首月尝鲜低至 9.4 元),开通后获得 `AWK_API_KEY`(主力模型 GLM-5.2 走此通道) +2. 注册 [SiliconFlow](https://cloud.siliconflow.cn/i/WNLYbBpi)(🎁 欢迎使用wiseflow邀请链接,注册认证后获得 16 元代金券),获得 `SILICONFLOW_API_KEY`(视觉/替补模型,必申请) > 如果习惯使用 ChatGPT / Gemini / Claude 等海外模型见下方[模型费用说明](#-模型费用说明)中的 AiHubMix 备选方案。 +> 🎬 **想用 5.5.2 的视频生成能力?** 需额外开通火山方舟 `doubao-seedance-2.0` 系列或阿里云百炼 `happyhorse-1.1` 系列模型,并把对应 key(`AWK_GEN_KEY` 或 `MODELSTUDIO_API_KEY`)配置到 `daemon.env`。详见下方[视频生成模型配置](#-视频生成模型配置)。 + ### 1. 获取代码 至 [Releases](https://github.com/TeamWiseFlow/wiseflow/releases) 下载最新版压缩包并解压; @@ -60,7 +64,7 @@ cd wiseflow - 拉取最新代码 - 初始化 `openclaw.json`(内置最佳模型配置,无需手动编辑) - 安装系统 daemon(开机自启 + 崩溃重启) -- **交互式引导你输入** `DEEPSEEK_API_KEY` 和 `SILICONFLOW_API_KEY`(仅在首次或缺失时询问) +- **交互式引导你输入** `AWK_API_KEY` 和 `SILICONFLOW_API_KEY`(仅在首次或缺失时询问) - 安装腾讯官方 `openclaw-weixin` extension,并引导扫码绑定 > **调试模式**(单次启动,适合测试):`./scripts/dev.sh gateway` @@ -99,14 +103,26 @@ cd wiseflow > > wiseflow5.x 底层基于 openclaw,Agent 工作流对 token 消耗有一定要求,建议先准备好大模型 API: > -> - **主力模型(强烈推荐)**:[DeepSeek 官方 API](https://platform.deepseek.com/) — 综合性能、速度、性价比最优。高缓存命中机制,实际应用成本可控。需要注册并充值获得 `DEEPSEEK_API_KEY`。 -> - **替补 & 视觉模型**:[SiliconFlow](https://cloud.siliconflow.cn/i/WNLYbBpi) — 模型丰富,可作为 DeepSeek 的 fallback,同时提供视觉理解模型(`Qwen/Qwen3.6-27B`)和生图/生视频 API。需要注册获得 `SILICONFLOW_API_KEY`。 +> - **主力模型(强烈推荐)**:[火山引擎方舟 Coding Plan](https://volcengine.com/L/dx-wt80li-I/) — 一个套餐覆盖 GLM-5.2、Kimi-K2.7、MiniMax-M3、DeepSeek-V4 系列、Doubao-Seed-2.0 系列等主流模型,**工具不限**,wiseflow 默认主力模型 GLM-5.2 即走此通道。需要注册并开通 Coding Plan 获得 `AWK_API_KEY`。 +> > 🎁 **通过 wiseflow 邀请链接** [https://volcengine.com/L/dx-wt80li-I/](https://volcengine.com/L/dx-wt80li-I/) **订阅**(邀请码 `5Y5A6L86`),可叠加 **9.5 折**优惠,首月尝鲜低至 **9.4 元**,订得越多折扣越大。 +> - **替补 & 视觉模型**:[SiliconFlow](https://cloud.siliconflow.cn/i/WNLYbBpi) — 模型丰富,可作为主力模型的 fallback,同时提供视觉理解模型(`Qwen/Qwen3.6-27B`)和生图/生视频 API。需要注册获得 `SILICONFLOW_API_KEY`。 > > 🎁 以上 SiliconFlow 链接为 wiseflow 邀请链接,通过此链接注册,你和 wiseflow 项目各可获得一张 16 元代金券。 > > - **海外模型用户**:如果想使用 ChatGPT / Gemini / Claude 等海外模型,可通过 [AiHubMix](https://aihubmix.com/?aff=Gp54) 统一接入(全兼容 OpenAI 接口,国内直连)。欢迎通过此[邀请链接](https://aihubmix.com/?aff=Gp54)注册。备选配置模板见 `config-templates/openclaw-aihubmix.json`。 > > 配置模板已预置以上最佳实践,`install.sh` 会自动检测所需环境变量并引导你输入。安装后重启 openclaw gateway 即可生效。 +> **🎬 视频生成模型配置** +> +> 5.5.2 的短视频制作(`video-product`)需额外开通视频生成模型,并把对应 key 配置到 `daemon.env`(任选其一,百炼优先): +> +> | 平台 | 环境变量 | 模型 | +> |------|---------|------| +> | 阿里云百炼(优先) | `MODELSTUDIO_API_KEY`(或 `DASHSCOPE_API_KEY`) | `happyhorse-1.1-i2v` / `happyhorse-1.1-t2v` / `happyhorse-1.1-r2v` | +> | 火山引擎方舟 | `AWK_GEN_KEY` | `doubao-seedance-2-0-fast-260128` / `doubao-seedance-2-0-260128` / `doubao-seedance-2-0-mini-260615` | +> +> 两个 key 都配了走百炼,只配 `AWK_GEN_KEY` 走火山,都没配则 `video-product` 自动降级为 pexels/pixabay 免费素材模式。注意 `AWK_GEN_KEY` 与主力模型的 `AWK_API_KEY` 是一个 key,但必须在环境变量中以不同变量名称赋值,火山视频生成只认 `AWK_GEN_KEY`。申请成功后可以让系统自带的全局IT Engineer帮你完成配置。 + 🎉 wiseflow 目前提供付费知识库,包含《手把手从零开始安装教程》、《安装之后三分钟上手指南》、《Openclaw自定义配置全案教程》、《Windows 下安装 WSL2 无脑教程》、《秘籍:云服务器(ECS)部署》以及各种最佳实践分享,年费仅需¥168,还能加入 **vip微信交流群** ,共同探讨交流各种玩法,还有每月一次的闭门分享(腾讯会议),陪伴你从“小白“到“大神“! 欢迎添加”掌柜的“企业微信(这背后接的就是 wiseflow sales-cs)咨询了解: @@ -133,20 +149,19 @@ Wiseflow 的做法是提供 `Crew Template`,针对每个岗位提供专属 ski | IT Engineer(默认启用,全局唯一,可协助其他 crew 排障) | 系统运维、配置、故障排查 | seo、icp-filing、icp-exemption、tccli、alicloud-find-skills、session-logs | | HRBP | 招募管理对外 crew,周期扫描 feedback 升级 | crew-recruit、crew-modify、crew-remove、crew-list、crew-usage | | 商务拓展 | 客户挖掘端 | 社交媒体潜客挖掘、竞品监控、行业情报、生成业务介绍 PPT | -| 自媒体运营 | 内容生产端 | 写稿、生图、15 个平台自动发布、t2video(短视频)、highlight-clipper(高光剪辑) | +| 自媒体运营 | 内容生产端 | 写稿、生图、短视频制作(`video-product`:多输入源 → 自动脚本 → 分段生成 → 合成)、视频分发(微信视频号/小红书/抖音/Twitter 等)、`de-mouth`(去口误)、`highlight-clipper`(高光剪辑)、15 个平台自动发布 | | 设计师 | 视觉设计端 | 15 套品牌设计系统、完整网页/APP/品牌视觉体系构建 | | 销售型客服 | 获客转化端 | 自动回复促进成交、调研用户来源、记录客户信息、发起/确认收款 | -| 投资人关系 | 融资端 | 寻找投资人、冷接触、填报申请表、生成 BP | -| Video Producer* | 专业视频端 | 专业短视频制作 | +| 投资人关系(预发布) | 融资端 | 寻找投资人、冷接触、填报申请表、生成 BP | -> *\* 标记的 crew 由 Pro 版本提供*
- 自媒体运营 — 支持的社交媒体平台(15 个) + 自媒体运营 — 支持的社交媒体平台(16 个) | 平台 | 发布方式 | |------|---------| | 微信公众号 | API + wenyan-cli 渲染 | + | 微信视频号* | 浏览器 + CDP(wujie shadow DOM) | | 企业微信朋友圈* | API | | 小红书* | API | | 抖音 | API(OAuth2) | @@ -230,11 +245,12 @@ wiseflow 通过 `patches/` 目录对 openclaw 源码打补丁,每次运行 `ap | 补丁 | 说明 | 相关环境变量 | |------|------|-------------| -| `001-relax-exec-allowlist-shell-syntax` | 放宽 exec-approvals 的 shell 语法限制,允许 `$()`、反引号、重定向等常用写法 | 无 | | `002-disable-web-search-env-var` | 支持通过环境变量禁用 openclaw 内置 web search | `OPENCLAW_DISABLE_WEB_SEARCH=1` | | `003-act-field-validation` | 修复浏览器 act 动作的字段验证逻辑 | 无 | -| `004-chrome-port-grace-retry` | Chrome CDP 端口占用时优雅重试,避免因端口冲突导致浏览器启动失败 | 无 | | `005-browser-timeout-env-var` | 支持通过环境变量自定义浏览器操作默认超时(原默认仅 20 秒,网络慢时容易中断) | `OPENCLAW_BROWSER_TIMEOUT_MS=60000` (执行 install.sh 脚本会自动配置)| +| `006-connectovercdp-no-defaults` | `connectOverCDP` 启用 `noDefaults: true`,避免 Patchright 修改用户浏览器状态 | 无 | + +> 旧补丁 `001-relax-exec-allowlist-shell-syntax`、`004-chrome-port-grace-retry` 已于 v5.5.2 删除:前者上游 exec 审批重构后复合命令已原生支持逐段匹配 allowlist,后者上游新增主动杀陈旧 Chrome 占用进程的恢复逻辑,均被上游吸收。详见 [CHANGELOG.md](CHANGELOG.md)。 #### 浏览器增强 @@ -325,6 +341,8 @@ wiseflow/ - 文颜(Markdown文章排版美化工具,支持微信公众号、今日头条、知乎等平台。) https://github.com/caol64/wenyan - Everything Claude Code(Claude Code 全局 skill / rule / agent 集合,wiseflow 的 complex-task 等编排 skill 借鉴了其 blueprint 和 gan-style-harness 的设计思路) https://github.com/affaan-m/everything-claude-code - awesome-design-md(A curated collection of design systems in markdown format — Designer 内置设计系统库参考了此项目的设计系统结构) https://github.com/VoltAgent/awesome-design-md +- videocut-skills(视频去口误/精剪技能集 — `de-mouth` 技能原汁原味借鉴其口误检测与剪映草稿生成能力) https://github.com/Ceeon/videocut-skills +- cheat-on-content(自媒体打分算法借鉴) https://github.com/XBuilderLAB/cheat-on-content ## Citation diff --git a/addons/README.md b/addons/README.md index fc1e8177..a1e383c7 100644 --- a/addons/README.md +++ b/addons/README.md @@ -7,7 +7,7 @@ This directory's subdirectories are **git-ignored** — third-party addons are n wiseflow 采用两级扩展机制: -- **Base wiseflow**(`patches/` + `skills/`):每次 `apply-addons.sh` 运行时无条件应用,对所有 addon 和 crew 生效。包括代码补丁(`patches/*.patch`)和默认全局技能(`skills/`)。 +- **Base wiseflow**(`patches/` + `skills/`):每次 `apply-addons.sh` 运行时无条件应用,对所有 addon 和 crew 生效。包括代码补丁(`patches/*.patch`)、插件(`patches/suppress-stale-reply`)和默认全局技能(`skills/`)。 - **Addon**(`addons/*/`):在 base 之上叠加,提供额外全局技能(`skills/`)和 Crew 模板(`crew/`)。 > **注意**:addon 不包含 patches 层。如需对 openclaw 打补丁,请将 patch 放到项目根目录的 `patches/` 下,而非 addon 内部。 diff --git a/addons/officials/addon.json b/addons/officials/addon.json index 04a743ed..d8fad313 100644 --- a/addons/officials/addon.json +++ b/addons/officials/addon.json @@ -2,8 +2,8 @@ "name": "wiseflow officials", "version": "0.5.0", "description": "官方 Crew 模板(selfmedia-operator / business-developer / designer / ir / sales-cs)+ 专属全局技能(rss-reader / siliconflow-img-gen / pexels-footage / pixabay-footage / council / connections-optimizer / email-ops / pitch-deck / ppt-maker / social-graph-ranker / web-form-fill / xhs-interact / xianyu-ops)", - "openclaw_version": "2026.5.28", - "openclaw_commit": "e93216080aa1f425d3ab127014603eba8e365b2d", + "openclaw_version": "2026.6.10", + "openclaw_commit": "aa69b12d0086b631b139c1435c9621a5783e3a40", "auto-activate": false, "internal_crews": ["business-developer", "designer", "ir", "selfmedia-operator"], "external_crews": ["sales-cs"] diff --git a/addons/officials/crew/business-developer/ALLOWED_COMMANDS b/addons/officials/crew/business-developer/ALLOWED_COMMANDS index 064cfcd2..dbcd4906 100644 --- a/addons/officials/crew/business-developer/ALLOWED_COMMANDS +++ b/addons/officials/crew/business-developer/ALLOWED_COMMANDS @@ -8,3 +8,12 @@ +./skills/info-record/scripts/record-content.sh +./skills/info-record/scripts/query-today.sh +sqlite3 ++nano-pdf ++jq ++rg ++tmux ++curl ++summarize ++gifgrep ++node ++python3 \ No newline at end of file diff --git a/addons/officials/crew/designer/USER.md b/addons/officials/crew/designer/USER.md index 7ef044c8..64c5814e 100644 --- a/addons/officials/crew/designer/USER.md +++ b/addons/officials/crew/designer/USER.md @@ -11,4 +11,4 @@ The user is the boss. - 用户大多数时候有明确的功能需求,但视觉语言表达不精确 - 用户对品牌规范可能不熟悉,需要设计师主动查阅 MEMORY.md 并提醒约束 - 用户希望一次看到多个方案选择,而不是只看一个版本 -- 用户可能会用"再改一下"这样的模糊反馈,需要主动追问具体不满意点 \ No newline at end of file +- 用户可能会用"再改一下"这样的模糊反馈,需要主动追问具体不满意点 diff --git a/addons/officials/crew/ir/ALLOWED_COMMANDS b/addons/officials/crew/ir/ALLOWED_COMMANDS index 70c73802..95357eaf 100644 --- a/addons/officials/crew/ir/ALLOWED_COMMANDS +++ b/addons/officials/crew/ir/ALLOWED_COMMANDS @@ -10,3 +10,12 @@ +./skills/ir-record/scripts/record-application.sh +./skills/ir-record/scripts/query-applications.sh +sqlite3 ++nano-pdf ++jq ++rg ++tmux ++curl ++summarize ++gifgrep ++node ++python3 diff --git a/addons/officials/crew/ir/DENIED_SKILLS b/addons/officials/crew/ir/DENIED_SKILLS index 52d370bd..ef9bdbd4 100644 --- a/addons/officials/crew/ir/DENIED_SKILLS +++ b/addons/officials/crew/ir/DENIED_SKILLS @@ -1,4 +1,3 @@ -github gh-issues coding-agent diff --git a/addons/officials/crew/sales-cs/ALLOWED_COMMANDS b/addons/officials/crew/sales-cs/ALLOWED_COMMANDS index 35ea3ec7..b0026dd8 100644 --- a/addons/officials/crew/sales-cs/ALLOWED_COMMANDS +++ b/addons/officials/crew/sales-cs/ALLOWED_COMMANDS @@ -1,7 +1,6 @@ # customer-service ALLOWED_COMMANDS # 在 T0 基础上精确放行声明式技能所需脚本 # 格式:+ 追加允许(相对于 workspace 根目录) - # customer-db 具名操作脚本(无原子 SQL 访问权限) +./skills/customer-db/scripts/cs-update.sh +./skills/customer-db/scripts/follow-up-create.sh @@ -10,6 +9,11 @@ +./skills/customer-db/scripts/follow-up-mark-sent.sh +./skills/customer-db/scripts/follow-up-complete.sh +./skills/customer-db/scripts/follow-up-expire.sh - +./skills/exp_invite/scripts/invite.sh +./skills/proactive-send/scripts/send.sh ++nano-pdf ++./skills/payment_confirm/scripts/send-subs.sh ++./skills/payment_confirm/scripts/confirm-payment.sh ++./skills/payment_confirm/scripts/send-kb.sh ++jq ++rg \ No newline at end of file diff --git a/addons/officials/crew/sales-cs/skills/exp_invite/SKILL.md b/addons/officials/crew/sales-cs/skills/exp_invite/SKILL.md index c9ef0364..eb28cd8d 100644 --- a/addons/officials/crew/sales-cs/skills/exp_invite/SKILL.md +++ b/addons/officials/crew/sales-cs/skills/exp_invite/SKILL.md @@ -25,23 +25,40 @@ description: > - `--user-id-external`:来自消息上下文 Sender 块的 `id` 字段(awada 原始用户 ID),用于 awada 平台路由邀请动作 ## 行为规则 -- 邀请消息不是发给用户看的自然语言,而是 awada 控制消息: + +### 主动邀请限制 +对于 `business_status` 不为 `free` 的用户(如 `exp_invited`、`subs`、`club`): +- **不要主动邀请**他们进入体验群 +- 应回到主流程 3.7,继续主动引导 + +### 允许再次邀请的情况 +如果客户**明确要求**再次拉入群或再次 invite,可以使用 `--force` 参数强制发送邀请: +- 客户可能忘记接受之前的邀请 +- 客户可能主动退出后想重新加入 + +```bash +./skills/exp_invite/scripts/invite.sh \ + --peer "<[CustomerDB].peer>" \ + --user-id-external "" \ + --force +``` + +### business_status 更新规则 +- 若 `business_status` 为空或 `free`:更新为 `exp_invited` +- 若已有值(`exp_invited`/`subs`/`club`)且使用 `--force`:**不更新** business_status,仅发送邀请 + +### 邀请消息格式 +邀请消息是 awada 控制消息,不是自然语言: ```text /invite////风暴眼(wiseflow情报小站) ``` -- awada-channel 会将其转为拉群动作 -- 发送前先查询数据库: - - 若当前 `business_status` 已是 `exp_invited`,则**不要重复邀请** - - 此时应回到主流程 3.7,继续主动引导 -- 若尚未邀请,则: - 1. 更新数据库中的 `business_status = exp_invited` - 2. 输出 invite 控制消息 +awada-channel 会将其转为拉群动作。 ## 返回约定 - 成功:标准输出 invite 控制消息 -- 已邀请过:输出 `ALREADY_INVITED`,并以非 0 状态退出 +- 已邀请过且未指定 --force:输出 `ALREADY_INVITED`,并以非 0 状态退出 ## 当前体验群名称 - `风暴眼(wiseflow情报小站)` diff --git a/addons/officials/crew/sales-cs/skills/exp_invite/scripts/invite.sh b/addons/officials/crew/sales-cs/skills/exp_invite/scripts/invite.sh index 0c231784..c7ee8361 100755 --- a/addons/officials/crew/sales-cs/skills/exp_invite/scripts/invite.sh +++ b/addons/officials/crew/sales-cs/skills/exp_invite/scripts/invite.sh @@ -2,11 +2,15 @@ # Send awada invite control message and update customer status to exp_invited. # --peer: DB primary key (from [CustomerDB].peer), used for all DB operations. # --user-id-external: raw awada user ID (from Sender.id), used for the invite routing message. +# --force: force invite even if business_status is not 'free' (for re-invite requests). set -euo pipefail +DB_FILE="./db/customer.db" + PEER="" USER_ID_EXTERNAL="" GROUP_NAME="风暴眼(wiseflow情报小站)" +FORCE="" while [ $# -gt 0 ]; do case "$1" in @@ -22,6 +26,10 @@ while [ $# -gt 0 ]; do GROUP_NAME="${2:-}" shift 2 ;; + --force) + FORCE="1" + shift + ;; *) echo "Unknown argument: $1" >&2 exit 1 @@ -42,19 +50,32 @@ fi WORKDIR="$(cd "$(dirname "$0")/../../.." && pwd)" cd "$WORKDIR" -./skills/customer-db/scripts/db.sh ensure >/dev/null +if [ ! -f "$DB_FILE" ]; then + echo "❌ Database not found: $DB_FILE" >&2 + exit 1 +fi -existing_status="$(./skills/customer-db/scripts/db.sh sql "SELECT business_status FROM cs_record WHERE peer = '$PEER'" | tail -n +2 | head -n 1 || true)" +sql_quote() { + printf '%s' "$1" | sed "s/'/''/g" +} + +existing_status="$(sqlite3 "$DB_FILE" "SELECT business_status FROM cs_record WHERE peer = '$(sql_quote "$PEER")'" || true)" if [ -z "$existing_status" ]; then - ./skills/customer-db/scripts/db.sh sql "INSERT INTO cs_record (peer, business_status, purpose, prompt_source) VALUES ('$PEER', 'free', '', '')" >/dev/null + sqlite3 "$DB_FILE" "INSERT INTO cs_record (peer, business_status, purpose, prompt_source) VALUES ('$(sql_quote "$PEER")', 'free', '', '')" existing_status="free" fi -if [ "$existing_status" = "exp_invited" ]; then +# Block auto-invite for non-free users, unless --force is specified +if [ "$existing_status" != "free" ] && [ -z "$FORCE" ]; then echo "ALREADY_INVITED" exit 10 fi -./skills/customer-db/scripts/db.sh sql "UPDATE cs_record SET business_status = 'exp_invited' WHERE peer = '$PEER'" >/dev/null +# Only update business_status to exp_invited if current status is free or empty +# For exp_invited/subs/club users with --force, don't change business_status +if [ "$existing_status" = "free" ] || [ -z "$existing_status" ]; then + sqlite3 "$DB_FILE" "UPDATE cs_record SET business_status = 'exp_invited' WHERE peer = '$(sql_quote "$PEER")'" +fi + printf '/invite//%s//%s\n' "$USER_ID_EXTERNAL" "$GROUP_NAME" diff --git a/addons/officials/crew/selfmedia-operator/AGENTS.md b/addons/officials/crew/selfmedia-operator/AGENTS.md index 1dafb1a9..636985b4 100644 --- a/addons/officials/crew/selfmedia-operator/AGENTS.md +++ b/addons/officials/crew/selfmedia-operator/AGENTS.md @@ -2,7 +2,7 @@ ## 素材积累 -素材积累来源包括:用户分享的飞书文档/网页链接、网络搜集、调用 skills 生成。 +素材积累来源包括:用户分享的飞书文档/网页链接、网络搜集、媒体文件等,或按用户要求使用相应技能生成的媒体文件。 **注意**:用户也可能时不时的通过私聊渠道分享一些要点、思路以及注意事项等,这些应该记在长期记忆 **MEMORY.md** 中。 @@ -18,51 +18,76 @@ index.md 格式为: - 来源:仅适用于用户分享和网络搜集 - prompt:仅适用于 skill 生成 -## 自媒体内容产出策略 +### 微信公众号内容对标 -### 核心原则 +如果用户提供了微信公众号账号或者微信公众号文章链接("https://mp.weixin.qq.com/"开头),可以使用 `generate-wenyan-theme` 技能,参考用户提供的账号或公众号文章,创建相似的公众号排版模板 -每篇文章或视频内容中都必须自然而合理的包含结合 wiseflow 的内容。生产前先回顾下 `MEMORY.md` +### 小红书内容对标 -### 文章(图文)内容生产通用约束 +如果用户需要对小红书图文内容进行对标,可以使用`xhs-content-ops`技能 -为每篇文章在 `output_articles/` 下创建独立文件夹作为工作区,结构如下: +## 自媒体内容产出 + +### 文章(图文)内容生产 + +用户会给出一个主题或写作思路,同时可能给出相关的参考资料(一段话、参考文章、图、视频等)。 + +这种情况下需要先为每篇文章在 `output_articles/` 下创建独立文件夹作为工作区,结构如下: ``` output_articles/ └── / # 文章英文题目作为文件夹名 - ├── article.md # 文章正文 + ├── article.md # 文章正文(按用户要求,结合用户给的资料书写) ├── cover.jpg # 封面图(必须) ├── img1.jpg # 配图1 ├── img2.jpg # 配图2 └── ... ``` -每篇文章都要有配图,包括封面图和正文配图 - **配图要求**: - -配图类型优先级: - - 1. **素材图**:日常积累的素材图,尤其是用户分享的 +- 每篇文章都要有配图,包括封面图和正文配图 +- 配图类型优先级: + - 1. 用户提供的素材。 + - 2. **素材图**:日常积累的素材图,尤其是用户分享的 - 存放在 `campaign_assets/` 目录 - - 直观展示 wiseflow 的能力和实际效果 - - 2. **技能生成图片**: + - 3. **技能生成图片**: - 优先使用 siliconflow-img-gen 生成,siliconflow-img-gen 不可用时,尝试 pexels-footage 或 pixabay-footage 下载免版权图片 -## Publish Strategy(发布执行策略) +按需写作的文章生产后**主动询问用户是否需要打分流程**。后续按用户决策推进(每一步决策由用户做): + +> 打分脚本(`score-only.sh`/`cal-toggle.sh`)与盲打分规范来自 `content-calibrator` 技能,发布则依据各个平台发布技能。 + +1. **问是否打分**。 + - 用户说**要打分** → 对 `article.md` 执行打分评估:主 agent `sessions_spawn` blind sub-agent(只喂 `article.md` + `calibration//rubric_notes.md`,输出 7 维分)→ 使用 `score-only.sh` 校验 + 判阈值门。平台未启用 calibration → 跳过打分并告知用户。 + - 每轮打分后,**询问用户是否发布**。 + - 用户有意见,则按用户意见修改之后再次执行打分流程,直到用户确认可发布。 + - 用户说**发布** → 调对应发布技能发布 + - 用户说**不必打分直接发布** → 直接调发布技能发布。 +2. 发布到哪个平台、是否多平台,由用户指定,因为涉及到用户交互和浏览器操作,所以多平台发布必须串行执行。 +3. 打分阈值取自 `calibration//.cheat-state.json` 的 `score_threshold`(默认 0=不拦截),每维需 > 阈值。打分流程与阈值命令见 `content-calibrator/SKILL.md`。 + +### 视频内容生产 + +**视频制作统一使用 `video-product` 技能**,它在 `video_generate` 工具基础上提供了素材获取、脚本编写、用户确认、合成组装、封面图制作等作业流程的指导,必须严格遵守。 + +支持按如下四种输入制作视频: +1. 文章链接(网页URL、本地文件、微信公众号文章) +2. 文字主题(用户直接给出主题或写作思路) +3. 用户已有素材(视频文件、图片参考) + +### 视频剪辑加工 + +你目前拥有两种简单的视频剪辑加工能力,适用于用户提供了原始视频素材,需要你帮忙进行剪辑的情况。 -### 文末统一宣传 hook +- `de-mouth` 技能用于处理口播视频,自动识别并删除静音、语气词、卡顿词、重复句、残句等,输出干净视频+字幕+剪映草稿。 +- `highlight-clipper` 技能用于自动从本地视频中提取高光片段。通过 ASR 转录 + 文本分析识别高光时刻,剪辑输出多段短视频。 -所有对外发布文章(视频则为简介区),必须按如下平台策略在文末添加统一宣传hook: +### 视频发布流程 - +> 打分脚本与盲打分规范来自 `content-calibrator` 技能,发布则依据各个平台发布技能。 -### 发布记录管理 +当用户确认视频制作内容后。先参考 `output_videos//scripts.md` 草拟视频发布的题目和简介以及hashtag。视频简介中应提及Wiselow,但不要有明显引流信息,更加禁止放二维码、联系方式等 -**统一使用 `published-track` 技能管理所有发布记录**。 +拟好后分别创建subagent(self-spawn)按用户指定发布的平台调用对应技能进行发布。但是对于使用浏览器自动化进行发布的技能(`twitter-post`, `wechat-channels-publish`)不可并行进行,避免浏览器资源竞态。 -- 数据库位置:`./db/published_track.db`(初始化:`./skills/published-track/scripts/init-db.sh`,幂等可重复执行) -- 按平台分表,每张表包含标题、类型、原始文件夹、发布 URL、发布日期、互动指标等字段 -- 所有发布技能在发布成功后必须自动调用 `record.sh` 记录 -- 数据更新通过 `update-metrics.sh` 完成(心跳巡检时使用) -- 查询通过 `query.sh` 和 `check-published.sh` 完成 \ No newline at end of file +你要负责跟进各个subagent的进展,避免他们长时间卡住,有问题及时反馈。如果某一个平台缺乏登录的credentials,或者浏览器缺乏登录态,及时反馈用户,让用户提供。用户提供后,你要按技能要求存储下来,以便后续使用。 diff --git a/addons/officials/crew/selfmedia-operator/ALLOWED_COMMANDS b/addons/officials/crew/selfmedia-operator/ALLOWED_COMMANDS index 67d59de9..06cabca2 100644 --- a/addons/officials/crew/selfmedia-operator/ALLOWED_COMMANDS +++ b/addons/officials/crew/selfmedia-operator/ALLOWED_COMMANDS @@ -1,3 +1,38 @@ # T2 基线已含 python3/node/npx/ffmpeg/sed/curl,无需重复声明 +cut +base64 ++nano-pdf ++jq ++rg ++tmux ++curl ++summarize ++gifgrep ++node ++python3 +# published-track ++./skills/published-track/scripts/init-db.sh ++./skills/published-track/scripts/record.sh ++./skills/published-track/scripts/update-metrics.sh ++./skills/published-track/scripts/query.sh ++./skills/published-track/scripts/check-published.sh ++./skills/published-track/scripts/query-pending.sh ++./skills/published-track/scripts/set-distribute-status.sh ++./skills/published-track/scripts/cal-toggle.sh ++./skills/published-track/scripts/score-only.sh ++./skills/published-track/scripts/migrate-v2.sh ++./skills/published-track/scripts/fetch-and-update-metrics.sh +# content-calibrator ++./skills/content-calibrator/scripts/init.sh ++./skills/content-calibrator/scripts/score-and-record.sh ++./skills/content-calibrator/scripts/query-metrics.sh ++./skills/content-calibrator/scripts/build-calibration-pool.sh ++./skills/content-calibrator/scripts/import-viral-chaser.sh +# wx-mp-publisher ++./skills/wx-mp-publisher/scripts/publish-wx-mp.sh +# de-mouth ++./skills/de-mouth/scripts/de_mouth.py +# viral-chaser ++./skills/viral-chaser/scripts/viral_chaser.sh +# xhs-content-ops ++./skills/xhs-content-ops/scripts/fetch_note_content.sh diff --git a/addons/officials/crew/selfmedia-operator/BUILTIN_SKILLS b/addons/officials/crew/selfmedia-operator/BUILTIN_SKILLS index a198c3e6..4e4a3727 100644 --- a/addons/officials/crew/selfmedia-operator/BUILTIN_SKILLS +++ b/addons/officials/crew/selfmedia-operator/BUILTIN_SKILLS @@ -1,5 +1,6 @@ wx-mp-publisher twitter-post +twitter-interact douyin-publish youtube-publish tiktok-publish @@ -7,13 +8,15 @@ facebook-publish instagram-publish threads-publish pinterest-publish -kuaishou-publish bilibili-publish wxwork-moments -xhs-content-ops xhs-interact +login-manager juejin-publish toutiao-publish pexels-footage pixabay-footage highlight-clipper +content-calibrator +de-mouth +video-product \ No newline at end of file diff --git a/addons/officials/crew/selfmedia-operator/HEARTBEAT.md b/addons/officials/crew/selfmedia-operator/HEARTBEAT.md index ccd7633a..5194396a 100644 --- a/addons/officials/crew/selfmedia-operator/HEARTBEAT.md +++ b/addons/officials/crew/selfmedia-operator/HEARTBEAT.md @@ -1,19 +1,19 @@ -# 心跳任务 +# 心跳/定时任务 -## 执行约束 +## 凌晨复盘任务 -1. **无时间限制**:HEARTBEAT/cron 触发后必须执行完清单全部内容 -2. **遇到技术故障时**: - - 先尝试关闭并重启浏览器 - - 仍不解决 → spawn IT Engineer 协助 - - 仍无法解决 → 跳过当前任务,继续后续步骤,不卡住整个流程 -3. **不可呼唤用户协助**(定时任务可能深夜执行) -4. **浏览器操作必须串行**,不可并行,避免竞态抢夺 +### 执行约束 ---- +1. **无时间限制**:任务执行不受深夜时间限制,必须执行完 HEARTBEAT 清单全部内容 -## 当前无定时任务 +2. **遇到技术故障时处理方案**: -如有任务需求,向用户了解清楚后,参照 `HEARTBEAT_TEMPLATE.md` 的格式写入对应工作模式配置。 + - 先尝试彻底关闭浏览器,再打开(使用默认 `openclaw` profile); + - 重启浏览器不解决问题时,**spawn IT Engineer**协助解决:调用 `sessions_spawn`,将问题现象、错误信息、当前任务上下文完整传递给 IT Engineer,请它协助解决; + - 仍无法解决 → **跳过当前任务,继续执行后续步骤**,不要卡住整个 HEARTBEAT -当前:回复 `HEARTBEAT_OK` + 不可: + - ❌ 呼唤用户协助解决,HEARTBEAT 在深夜执行,喊用户也没用 + - ❌ 不可中断任务,通过以上三步依然无法进行的任务则跳过,继续执行后续步骤,绝对不允许中断HEARTBEAT! + +3. HEARTBEAT 任务涉及大量浏览器操作,因此涉及浏览器的 subagent 任务不能并行发,必须排队串行执行,避免浏览器竞态抢夺。 diff --git a/addons/officials/crew/selfmedia-operator/IDENTITY.md b/addons/officials/crew/selfmedia-operator/IDENTITY.md index 1ceb0ce3..5b49047d 100644 --- a/addons/officials/crew/selfmedia-operator/IDENTITY.md +++ b/addons/officials/crew/selfmedia-operator/IDENTITY.md @@ -4,7 +4,7 @@ 小编 ## Role -业务驱动型自媒体内容专家 — 以推广公司产品与业务为核心目标,深耕主流自媒体生态,发现热点、采集素材、撰写图文、指导视频生产,交付可直接发布的内容,并深入自媒体平台上各个账号的运营。 +业务驱动型自媒体内容专家 — 以推广公司产品与业务为核心目标,深耕主流自媒体生态,发现热点、采集素材、撰写图文、视频生产,交付可直接发布的内容,并深入自媒体平台上各个账号的运营。 ## Personality 贴地气、有洞察力、执行力强。能感知平台气氛和受众喜好,把枯燥的信息变成有传播力的图文或视频。讲究效率,稿件出炉前必请用户确认。 diff --git a/addons/officials/crew/selfmedia-operator/SOUL.md b/addons/officials/crew/selfmedia-operator/SOUL.md index fb4f2a4d..1312287b 100644 --- a/addons/officials/crew/selfmedia-operator/SOUL.md +++ b/addons/officials/crew/selfmedia-operator/SOUL.md @@ -7,8 +7,6 @@ ## 公司与业务背景信息 - - ## Core Responsibilities ### 素材管理 diff --git a/addons/officials/crew/selfmedia-operator/TOOLS.md b/addons/officials/crew/selfmedia-operator/TOOLS.md index b443d4d2..4ad93a21 100644 --- a/addons/officials/crew/selfmedia-operator/TOOLS.md +++ b/addons/officials/crew/selfmedia-operator/TOOLS.md @@ -3,3 +3,25 @@ ## 环境备注 - 文生图/改图默认输出 JPG 格式:企业微信后台发送图片只支持 JPG;如需 PNG 需显式指定 --format png + +### 📝 视频封面/海报制作经验 + +`siliconflow-img-gen` 可以很好的直接出带文字的海报,完全不必要先生成图,然后自己再编写脚本拼字。 + +具体见 `siliconflow-img-gen` 技能中 `视频封面/海报最佳实践`。 + +但是**绝对不要用 image_generate 工具(默认 minimax/image-01)直接出带字图片!** + +实际测试下来`minimax/image-01`无法正确出带文字的图片。 + +### 数据库查询一定走 published-track 脚本 + +`sqlite3` 不在 allowlist 中。查询 published-track 数据库必须通过已有脚本: + +``` +✅ ./skills/published-track/scripts/query.sh --platform wx_mp +✅ ./skills/published-track/scripts/query-pending.sh + +❌ sqlite3 db/published_track.db "SELECT ..." +❌ echo ".tables" | sqlite3 db/published_track.db +``` \ No newline at end of file diff --git a/addons/officials/crew/selfmedia-operator/calibration/wx_mp/.cheat-state.json b/addons/officials/crew/selfmedia-operator/calibration/wx_mp/.cheat-state.json new file mode 100644 index 00000000..8fafc1ba --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/calibration/wx_mp/.cheat-state.json @@ -0,0 +1,20 @@ +{ + "schema_version": 2, + "platform": "wx_mp", + "mode": "cold-start", + "content_form": "长文", + "rubric_version": "v0", + "calibration_samples": 0, + "baseline_plays": null, + "typical_word_count": 2000, + "retro_window_days": 3, + "in_progress_session": null, + "pending_retros": [], + "consecutive_directional_errors": [], + "last_prediction_self_scored": false, + "last_bump_at": null, + "last_bump_self_audited": null, + "calibration_samples_at_last_bump": 0, + "enabled_perf_adapters": ["wx_mp"], + "created_at": "2026-06-14T00:00:00+08:00" +} diff --git a/addons/officials/crew/selfmedia-operator/calibration/wx_mp/audience.md b/addons/officials/crew/selfmedia-operator/calibration/wx_mp/audience.md new file mode 100644 index 00000000..f104853f --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/calibration/wx_mp/audience.md @@ -0,0 +1,15 @@ +# Audience — 受众画像 + +> 从复盘评论聚类派生。blind sub-agent **不可读**此文件。 + +--- + +## 基本画像 + +(复盘后从评论关键词聚类填充。) + +--- + +## 互动偏好 + +(哪些类型的内容获得更多互动?哪些评论模因反复出现?) diff --git a/addons/officials/crew/selfmedia-operator/calibration/wx_mp/benchmark.md b/addons/officials/crew/selfmedia-operator/calibration/wx_mp/benchmark.md new file mode 100644 index 00000000..06eb2986 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/calibration/wx_mp/benchmark.md @@ -0,0 +1,16 @@ +# Benchmark — 对标账号 + +> 导入对标账号后,记录对标信号和 pattern。 +> 由 content-calibrator 的 LearnFrom 操作维护。 + +--- + +## 对标账号列表 + +(暂无。运行"导入对标"添加。) + +--- + +## Pattern 提炼 + +(从对标内容中提取的结构 pattern,如开头方式、转折技巧、金句模式等。) diff --git a/addons/officials/crew/selfmedia-operator/calibration/wx_mp/rubric-memo.md b/addons/officials/crew/selfmedia-operator/calibration/wx_mp/rubric-memo.md new file mode 100644 index 00000000..7e6a41bd --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/calibration/wx_mp/rubric-memo.md @@ -0,0 +1,23 @@ +# Rubric Memo — 观察记录 + +> 本文件记录复盘产出的观察、实绩证据和样本引用。 +> **blind sub-agent 不读此文件**——它只读 rubric_notes.md。 +> rubric_notes.md 只放通用公式和维度定义,不含视频名/实绩/评论。 + +--- + +## 观察记录 + +(复盘后观察追加于此。每条观察必须可追溯到具体数据点。) + +--- + +## Benchmark 参考 + +(导入对标账号后,对标信号记录于此。) + +--- + +## Bump 升级 Memo + +(每次 rubric 升级后,append 升级详情含证据+诊断。) diff --git a/addons/officials/crew/selfmedia-operator/calibration/wx_mp/rubric_notes.md b/addons/officials/crew/selfmedia-operator/calibration/wx_mp/rubric_notes.md new file mode 100644 index 00000000..8fcadcfc --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/calibration/wx_mp/rubric_notes.md @@ -0,0 +1,72 @@ +# Rubric Notes — 评分公式 + +> **当前版本**: v0 +> **平台**: wx_mp(微信公众号) +> **内容形态**: 长文 +> **Last bumped at**: —(初始版本) +> **Upgrade memos**: 见 [rubric-memo.md](rubric-memo.md) + +--- + +## 当前评分维度 + +| 维度 | 代号 | 0 分 | 5 分 | 权重 | +|------|------|------|------|------| +| 情感共鸣 | ER | 纯信息罗列,无情感触点 | 读者强烈代入"说的就是我",有具象画面或经历 | ×1.5 | +| 钩子强度 | HP | 标题平庸,开头无悬念 | 标题/开头一句话锁定注意力,制造信息差或反差 | ×1.5 | +| 社会议题共振 | SR | 纯个人/产品向,无社会讨论 | 触及当下社会讨论,有立场可议 | ×1.5 | +| 金句密度 | QL | 全文无独立可传播的表达 | ≥3 句可脱离上下文独立传播的金句 | ×1.0 | +| 叙事性 | NA | 纯观点堆砌,无故事弧线 | 清晰的起承转合,读者被故事牵引 | ×1.0 | +| 受众广度 | AB | 极窄垂直,仅特定人群关心 | 跨人群普适(如搞钱、职场、AI焦虑) | ×1.0 | +| 实用价值 | PV | 纯情绪/观点,无可操作信息 | 读者可获得具体方法/工具/步骤 | ×1.0 | + +## 综合分公式 + +``` +composite = (ER×1.5 + HP×1.5 + SR×1.5 + QL + NA + AB + PV) / 8.5 × 2.0 +``` + +- 归一化常数: 8.5 +- 缩放因子: 2.0 +- 理论范围: 0 - 10 +- 整数维度分,composite 保留两位小数 + +## Bucket 方案(cold-start 等权占位) + +> cold-start 期 bucket 数字是 false precision,前 5 篇不给 bucket 概率分布。 +> 第 5 篇复盘后按实绩数据派生 bucket 边界。 + +| 档位 | 含义 | 边界(待校准) | +|------|------|---------------| +| 退步 | 低于基线 | < baseline × 0.3 | +| 持平 | 基线水平 | baseline × 0.3 ~ 1 | +| 命中 | 正常表现 | baseline × 1 ~ 3 | +| 小爆 | 超预期 | baseline × 3 ~ 10 | +| 大爆 | 现象级 | > baseline × 10 | + +--- + +## 版本速查 + +| 版本 | 公式签名 | 日期 | +|------|---------|------| +| v0 | ER1.5+HP1.5+SR1.5+QL+NA+AB+PV / 8.5×2 | 2026-06-14 | + +--- + +## 维度与权重变更规则 + +**维度和权重可以被修改,但必须满足以下条件之一**: +1. **用户主动要求** — "给公众号加个 XX 维度" / "把 SR 权重调到 2.0" +2. **Agent 提议 + 用户确认** — Agent 检测到系统性偏差后提议变更,必须等待用户明确同意才生效 + +变更流程: +- 变更维度(增/删/替换)→ 走 Bump 全量重打 + 排序一致性校验 +- 变更权重 → 走 Bump 流程 +- 变更被拒绝 → rubric 不动,观察记入 rubric-memo.md + +--- + +## 待验证假设 + +(复盘后观察会写入此处,bump 时验证或推翻) diff --git a/addons/officials/crew/selfmedia-operator/calibration/xhs/.cheat-state.json b/addons/officials/crew/selfmedia-operator/calibration/xhs/.cheat-state.json new file mode 100644 index 00000000..9e970b9c --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/calibration/xhs/.cheat-state.json @@ -0,0 +1,20 @@ +{ + "schema_version": 2, + "platform": "xhs", + "mode": "cold-start", + "content_form": "图文/视频笔记", + "rubric_version": "v0", + "calibration_samples": 0, + "baseline_plays": null, + "typical_word_count": 500, + "retro_window_days": 3, + "in_progress_session": null, + "pending_retros": [], + "consecutive_directional_errors": [], + "last_prediction_self_scored": false, + "last_bump_at": null, + "last_bump_self_audited": null, + "calibration_samples_at_last_bump": 0, + "enabled_perf_adapters": ["xhs"], + "created_at": "2026-06-14T00:00:00+08:00" +} diff --git a/addons/officials/crew/selfmedia-operator/calibration/xhs/audience.md b/addons/officials/crew/selfmedia-operator/calibration/xhs/audience.md new file mode 100644 index 00000000..f104853f --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/calibration/xhs/audience.md @@ -0,0 +1,15 @@ +# Audience — 受众画像 + +> 从复盘评论聚类派生。blind sub-agent **不可读**此文件。 + +--- + +## 基本画像 + +(复盘后从评论关键词聚类填充。) + +--- + +## 互动偏好 + +(哪些类型的内容获得更多互动?哪些评论模因反复出现?) diff --git a/addons/officials/crew/selfmedia-operator/calibration/xhs/benchmark.md b/addons/officials/crew/selfmedia-operator/calibration/xhs/benchmark.md new file mode 100644 index 00000000..28696b97 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/calibration/xhs/benchmark.md @@ -0,0 +1,16 @@ +# Benchmark — 对标账号 + +> 导入对标账号后,记录对标信号和 pattern。 +> 由 content-calibrator 的 LearnFrom 操作维护。 + +--- + +## 对标账号列表 + +(暂无。运行"导入对标 --platform xhs"添加。) + +--- + +## Pattern 提炼 + +(从对标内容中提取的结构 pattern,如封面风格、标题写法、话题标签策略、种草话术等。) diff --git a/addons/officials/crew/selfmedia-operator/calibration/xhs/rubric-memo.md b/addons/officials/crew/selfmedia-operator/calibration/xhs/rubric-memo.md new file mode 100644 index 00000000..26360a5b --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/calibration/xhs/rubric-memo.md @@ -0,0 +1,10 @@ +# Rubric Memo — 观察记录 + +> **blind sub-agent 硬禁读此文件**(含实绩数据,会污染盲打分) +> 由 Bump 和 Retro 操作维护。被推翻/吸收的观察删除,git history 是档案。 + +--- + +## 观察记录 + +(复盘后从实绩数据中提炼的观察会写入此处) diff --git a/addons/officials/crew/selfmedia-operator/calibration/xhs/rubric_notes.md b/addons/officials/crew/selfmedia-operator/calibration/xhs/rubric_notes.md new file mode 100644 index 00000000..2be7cb94 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/calibration/xhs/rubric_notes.md @@ -0,0 +1,72 @@ +# Rubric Notes — 评分公式 + +> **当前版本**: v0 +> **平台**: xhs(小红书) +> **内容形态**: 图文/视频笔记 +> **Last bumped at**: —(初始版本) +> **Upgrade memos**: 见 [rubric-memo.md](rubric-memo.md) + +--- + +## 当前评分维度 + +| 维度 | 代号 | 0 分 | 5 分 | 权重 | +|------|------|------|------|------| +| 情感共鸣 | ER | 纯信息罗列,无情感触点 | 读者强烈代入"说的就是我",有具象画面或经历 | ×1.5 | +| 钩子强度 | HP | 标题平庸,封面无吸引力 | 封面/标题一句话锁定注意力,制造信息差或反差 | ×1.5 | +| 社会议题共振 | SR | 纯个人/产品向,无社会讨论 | 触及当下社会讨论,有立场可议 | ×1.5 | +| 金句密度 | QL | 全文无独立可传播的表达 | ≥3 句可脱离上下文独立传播的金句 | ×1.0 | +| 叙事性 | NA | 纯观点堆砌,无故事弧线 | 清晰的起承转合,读者被故事牵引 | ×1.0 | +| 受众广度 | AB | 极窄垂直,仅特定人群关心 | 跨人群普适(如搞钱、职场、AI焦虑) | ×1.0 | +| 实用价值 | PV | 纯情绪/观点,无可操作信息 | 读者可获得具体方法/工具/步骤 | ×1.0 | + +## 综合分公式 + +``` +composite = (ER×1.5 + HP×1.5 + SR×1.5 + QL + NA + AB + PV) / 8.5 × 2.0 +``` + +- 归一化常数: 8.5 +- 缩放因子: 2.0 +- 理论范围: 0 - 10 +- 整数维度分,composite 保留两位小数 + +## Bucket 方案(cold-start 等权占位) + +> cold-start 期 bucket 数字是 false precision,前 5 篇不给 bucket 概率分布。 +> 第 5 篇复盘后按实绩数据派生 bucket 边界。 + +| 档位 | 含义 | 边界(待校准) | +|------|------|---------------| +| 退步 | 低于基线 | < baseline × 0.3 | +| 持平 | 基线水平 | baseline × 0.3 ~ 1 | +| 命中 | 正常表现 | baseline × 1 ~ 3 | +| 小爆 | 超预期 | baseline × 3 ~ 10 | +| 大爆 | 现象级 | > baseline × 10 | + +--- + +## 版本速查 + +| 版本 | 公式签名 | 日期 | +|------|---------|------| +| v0 | ER1.5+HP1.5+SR1.5+QL+NA+AB+PV / 8.5×2 | 2026-06-14 | + +--- + +## 维度与权重变更规则 + +**维度和权重可以被修改,但必须满足以下条件之一**: +1. **用户主动要求** — "给小红书加个 XX 维度" / "把 SR 权重调到 2.0" +2. **Agent 提议 + 用户确认** — Agent 检测到系统性偏差后提议变更,必须等待用户明确同意才生效 + +变更流程: +- 变更维度(增/删/替换)→ 走 Bump 全量重打 + 排序一致性校验 +- 变更权重 → 走 Bump 流程 +- 变更被拒绝 → rubric 不动,观察记入 rubric-memo.md + +--- + +## 待验证假设 + +(复盘后观察会写入此处,bump 时验证或推翻) diff --git a/addons/officials/crew/selfmedia-operator/openclaw_setting_sample.json b/addons/officials/crew/selfmedia-operator/openclaw_setting_sample.json index 211a30ae..35173fbf 100644 --- a/addons/officials/crew/selfmedia-operator/openclaw_setting_sample.json +++ b/addons/officials/crew/selfmedia-operator/openclaw_setting_sample.json @@ -4,7 +4,6 @@ "tiktok-post", "instagram-post", "youtube-upload", - "xhs-content-ops", "xhs-interact", "juejin-publish", "toutiao-publish", @@ -16,10 +15,10 @@ "wxwork-moments", "wxwork-drive", "wx-mp-publisher", - "complex-task" + "video-product" ], "subagents": { - "allowAgents": ["it-engineer", "designer", "video-producer"] + "allowAgents": ["it-engineer", "designer"] }, "maxConcurrent": 2, "tools": {} diff --git a/addons/officials/crew/selfmedia-operator/scripts/crop_watermarks.py b/addons/officials/crew/selfmedia-operator/scripts/crop_watermarks.py new file mode 100644 index 00000000..363be09c --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/scripts/crop_watermarks.py @@ -0,0 +1,37 @@ +#!/usr/bin/env python3 +""" +批量裁剪图片底部水印(知乎等平台右下角账号水印) +用法: python3 crop_watermarks.py <图片目录> +示例: python3 crop_watermarks.py ./output_articles/xxx/images +""" +from PIL import Image +import os +import sys + +if len(sys.argv) < 2: + print("用法: python3 crop_watermarks.py <图片目录>") + sys.exit(1) + +img_dir = sys.argv[1] +crop_h = 60 # 裁剪底部高度(px),覆盖知乎典型水印区域 + +if not os.path.isdir(img_dir): + print(f"错误: {img_dir} 不是有效目录") + sys.exit(1) + +for fname in sorted(os.listdir(img_dir)): + if not fname.endswith(('.jpg', '.png', '.jpeg', '.webp')): + continue + path = os.path.join(img_dir, fname) + img = Image.open(path) + w, h = img.size + if h > crop_h + 100: + cropped = img.crop((0, 0, w, h - crop_h)) + if cropped.mode != 'RGB': + cropped = cropped.convert('RGB') + cropped.save(path, 'JPEG', quality=92) + print(f"{fname}: {w}x{h} → {w}x{h - crop_h} (已裁剪底部 {crop_h}px)") + else: + print(f"{fname}: {w}x{h} 太小,跳过") + +print("完成") diff --git a/addons/officials/crew/selfmedia-operator/scripts/process_images.py b/addons/officials/crew/selfmedia-operator/scripts/process_images.py new file mode 100644 index 00000000..10f5820c --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/scripts/process_images.py @@ -0,0 +1,48 @@ +#!/usr/bin/env python3 +"""PIL post-process generated images: convert PNG to JPG, slight sharpening, copy to article folder.""" + +from PIL import Image, ImageEnhance, ImageFilter +import shutil +from pathlib import Path + +WORKDIR = Path("/home/wukong/.openclaw/workspace-media-operator") +ARTICLE_DIR = WORKDIR / "output_articles" / "freelancer-776-days" +ARTICLE_IMAGES = ARTICLE_DIR / "images" +ARTICLE_IMAGES.mkdir(parents=True, exist_ok=True) + +# (source_dir, target_name) +SOURCES = [ + (WORKDIR / "tmp" / "freelancer776_cover" / "00.png", "cover.jpg", "cover"), + (WORKDIR / "tmp" / "freelancer776_img1" / "00.png", "img1.jpg", "income"), + (WORKDIR / "tmp" / "freelancer776_img2" / "00.png", "img2.jpg", "time"), + (WORKDIR / "tmp" / "freelancer776_img3" / "00.png", "img3.jpg", "platforms"), + (WORKDIR / "tmp" / "freelancer776_img4" / "00.png", "img4.jpg", "milestone"), +] + +def post_process(src: Path, dst: Path) -> tuple[int, int]: + """Open PNG, apply slight sharpen, save as JPG quality 92.""" + img = Image.open(src).convert("RGB") + # Slight unsharp mask to clean AI softness + img = img.filter(ImageFilter.UnsharpMask(radius=1.2, percent=110, threshold=2)) + # Slight contrast boost + img = ImageEnhance.Contrast(img).enhance(1.05) + # Save as JPG + img.save(dst, "JPEG", quality=92, optimize=True) + w, h = img.size + return w, h + +print("=== Post-processing images ===\n") +for src, name, label in SOURCES: + if not src.exists(): + print(f" ❌ {label}: source missing {src}") + continue + # in-article images go to images/, cover to root + if name == "cover.jpg": + dst = ARTICLE_DIR / name + else: + dst = ARTICLE_IMAGES / name + w, h = post_process(src, dst) + size_kb = dst.stat().st_size / 1024 + print(f" ✅ {label}: {w}x{h}, {size_kb:.0f}KB → {dst.relative_to(WORKDIR)}") + +print("\n=== Done ===") diff --git a/addons/officials/crew/selfmedia-operator/skills/bilibili-publish/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/bilibili-publish/SKILL.md new file mode 100644 index 00000000..5ad840aa --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/bilibili-publish/SKILL.md @@ -0,0 +1,121 @@ +--- +name: bilibili-publish +description: Publish videos to Bilibili (B站) via Open Platform API (OAuth2). + Supports chunked video upload, cover image, tags, and partition selection. + Requires BILIBILI_APP_ID and BILIBILI_APP_SECRET environment variables. +metadata: + openclaw: + emoji: 📺 + requires: + bins: + - python3 +--- + +# B站视频发布(bilibili-publish) + +通过 B站开放平台 API(OAuth2 + HMAC-SHA256 签名)上传发布视频,支持分块上传、封面、分区和标签。 + +> ⚠️ 本技能使用 B站开放平台 OAuth2 API(`member.bilibili.com/arcopen/`),非 Web 创作中心 API。 +> 需要在 [B站开放平台](https://open.bilibili.com/) 申请 app_key/app_secret。 + +--- + +## 前置条件 + +1. 环境变量 `BILIBILI_APP_ID` 和 `BILIBILI_APP_SECRET` 已配置(B站开放平台凭据) +2. OAuth2 授权已完成(`~/.openclaw/logins/bilibili-oauth.json` 存在且有效) +3. 若 token 过期,脚本会自动尝试 refresh;若 refresh 也失败,需重新授权 + +--- + +## OAuth2 授权流程(首次使用) + +1. 获取授权页面 URL: + ``` + https://account.bilibili.com/pc/account-pc/auth/oauth?client_id=${BILIBILI_APP_ID}&gourl=${REDIRECT_URL}&state=random + ``` +2. 用户在浏览器中打开该 URL,扫码登录并授权 +3. 授权后回调 URL 中获取 `code` 参数 +4. 交换 token: + ```bash + python3 ./skills/bilibili-publish/scripts/publish_bilibili.py --exchange-token + ``` +5. Token 保存在 `~/.openclaw/logins/bilibili-oauth.json` + +--- + +## 使用方式 + +```bash +python3 ./skills/bilibili-publish/scripts/publish_bilibili.py \ + --title "视频标题" \ + --video video.mp4 \ + --tid 122 \ + --tags AI,科技,工具 +``` + +带封面和描述: + +```bash +python3 ./skills/bilibili-publish/scripts/publish_bilibili.py \ + --title "视频标题" \ + --video video.mp4 \ + --cover cover.jpg \ + --desc "视频描述" \ + --tid 122 \ + --tags AI,科技 +``` + +--- + +## 参数说明 + +| 参数 | 必填 | 说明 | +|------|------|------| +| `--title` | 是 | 视频标题,最多 80 字 | +| `--video` | 是 | 视频文件路径,支持 mp4 | +| `--cover` | 否 | 封面图路径;不提供则 B站自动截取 | +| `--desc` | 否 | 视频描述 | +| `--tid` | 否 | 分区 ID,默认 122(野生技术协会) | +| `--tags` | 是 | 逗号分隔的标签,最多 10 个,每个最多 20 字 | +| `--copyright` | 否 | 1=自制(默认),2=转载 | +| `--exchange-token` | 否 | OAuth2 授权码交换模式 | + +--- + +## 常用分区 ID + +| tid | 分区 | +|-----|------| +| 122 | 野生技术协会 | +| 36 | 知识 · 科技 | +| 95 | 数码 | +| 207 | 资讯 | +| 21 | 日常 | +| 76 | 美食制作 | + +--- + +## Agent 工作流 + +1. 确认 `BILIBILI_APP_ID` 和 `BILIBILI_APP_SECRET` 环境变量已设置 +2. 准备视频文件 + 标题 + 分区 + 标签 +3. 运行 `publish_bilibili.py` 脚本(自动完成上传+提交) +4. 检查 stdout JSON 输出: + - `{"ok": true, "bvid": "BVxxx", "url": "..."}` → 成功 + - `{"ok": false, "error": "CREDENTIALS_MISSING"}` → 需配置环境变量 + - `{"ok": false, "error": "AUTH_REQUIRED"}` → 需完成 OAuth2 授权流程 + - `{"ok": false, "error": "AUTH_EXPIRED"}` → 需重新授权 + - `{"ok": false, "error": "..."}` → 其他错误,反馈用户 + +--- + +## 错误处理 + +| 错误 | 原因 | 处理 | +|------|------|------| +| CREDENTIALS_MISSING | 环境变量未配置 | 设置 BILIBILI_APP_ID 和 BILIBILI_APP_SECRET | +| AUTH_REQUIRED | 无 OAuth token | 完成 OAuth2 授权流程 | +| AUTH_EXPIRED | token 过期且刷新失败 | 重新走授权流程获取新 code | +| UPLOAD_FAILED | 上传失败 | 检查网络和文件大小,重试一次 | +| SUBMIT_FAILED | 提交失败 | 检查分区和标签是否合法 | diff --git a/addons/officials/crew/selfmedia-operator/skills/bilibili-publish/scripts/publish_bilibili.py b/addons/officials/crew/selfmedia-operator/skills/bilibili-publish/scripts/publish_bilibili.py new file mode 100755 index 00000000..aa836c96 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/bilibili-publish/scripts/publish_bilibili.py @@ -0,0 +1,365 @@ +#!/usr/bin/env python3 +"""Publish videos to Bilibili via Open Platform API (OAuth2 + HMAC-SHA256). + +Based on AiToEarn's bilibili-api.service.ts, adapted for wiseflow skill architecture. + +Authentication: OAuth2 with app_key/app_secret (Bilibili Open Platform). + - BILIBILI_APP_ID and BILIBILI_APP_SECRET must be set in environment. + - Access token stored in ~/.openclaw/logins/bilibili-oauth.json + +Usage: + python3 publish_bilibili.py --title "标题" --video video.mp4 --tags AI,科技 +""" + +import argparse +import hashlib +import hmac +import json +import os +import sys +import time +import uuid +from pathlib import Path + +import requests + +LOGINS_DIR = Path.home() / ".openclaw" / "logins" +OAUTH_FILE = LOGINS_DIR / "bilibili-oauth.json" +CHUNK_SIZE = 5 * 1024 * 1024 # 5MB chunks + +# Open Platform API endpoints +AUTH_BASE = "https://api.bilibili.com/x/account-oauth2/v1" +MEMBER_BASE = "https://member.bilibili.com/arcopen/fn" +UPLOAD_BASE = "https://openupos.bilivideo.com/video/v2" + + +def output(data: dict) -> None: + sys.stdout.write(json.dumps(data, ensure_ascii=False) + "\n") + + +def err_exit(msg: str, code: int = 1) -> None: + sys.stderr.write(f"[bilibili-publish] ERROR: {msg}\n") + output({"ok": False, "error": msg}) + sys.exit(code) + + +# ── OAuth2 ──────────────────────────────────────────────────────────────── + +def get_app_credentials() -> tuple[str, str]: + app_id = os.environ.get("BILIBILI_APP_ID", "") + app_secret = os.environ.get("BILIBILI_APP_SECRET", "") + if not app_id or not app_secret: + err_exit( + "CREDENTIALS_MISSING: BILIBILI_APP_ID and BILIBILI_APP_SECRET " + "environment variables are required. Apply at https://open.bilibili.com/", + 2, + ) + return app_id, app_secret + + +def load_token() -> dict | None: + if not OAUTH_FILE.exists(): + return None + try: + return json.loads(OAUTH_FILE.read_text()) + except (json.JSONDecodeError, OSError): + return None + + +def save_token(data: dict) -> None: + LOGINS_DIR.mkdir(parents=True, exist_ok=True) + OAUTH_FILE.write_text(json.dumps(data, indent=2, ensure_ascii=False)) + + +def get_access_token() -> str: + """Get a valid access token, refreshing if needed.""" + app_id, app_secret = get_app_credentials() + token_data = load_token() + + if not token_data: + err_exit( + "AUTH_REQUIRED: No OAuth token found. Complete OAuth2 flow first:\n" + "1. Open auth page: login-manager or browser to get authorization code\n" + "2. Run: python3 publish_bilibili.py --exchange-token ", + 2, + ) + + access_token = token_data.get("access_token", "") + expires_at = token_data.get("expires_at", 0) + refresh_token = token_data.get("refresh_token", "") + + # Token still valid (with 10min buffer) + if access_token and expires_at > time.time() + 600: + return access_token + + # Try refresh + if refresh_token: + sys.stderr.write("[bilibili-publish] refreshing access token...\n") + resp = requests.post( + f"{AUTH_BASE}/refresh_token", + params={ + "client_id": app_id, + "client_secret": app_secret, + "grant_type": "refresh_token", + "refresh_token": refresh_token, + }, + timeout=30, + ) + data = resp.json() + if data.get("code") == 0 and data.get("data"): + token_info = data["data"] + save_token({ + "access_token": token_info["access_token"], + "refresh_token": token_info["refresh_token"], + "expires_at": int(time.time()) + token_info.get("expires_in", 2592000), + "mid": token_info.get("mid", ""), + }) + return token_info["access_token"] + sys.stderr.write(f"[bilibili-publish] token refresh failed: {data}\n") + + err_exit( + "AUTH_EXPIRED: Token expired and refresh failed. Re-authorization required.", + 2, + ) + + +def exchange_token(code: str) -> None: + """Exchange authorization code for access token.""" + app_id, app_secret = get_app_credentials() + resp = requests.post( + f"{AUTH_BASE}/token", + params={ + "client_id": app_id, + "client_secret": app_secret, + "grant_type": "authorization_code", + "code": code, + }, + timeout=30, + ) + data = resp.json() + if data.get("code") != 0: + err_exit(f"TOKEN_EXCHANGE_FAILED: {data}") + + token_info = data["data"] + save_token({ + "access_token": token_info["access_token"], + "refresh_token": token_info["refresh_token"], + "expires_at": int(time.time()) + token_info.get("expires_in", 2592000), + "mid": token_info.get("mid", ""), + }) + output({"ok": True, "mid": token_info.get("mid", ""), "path": str(OAUTH_FILE)}) + + +# ── HMAC-SHA256 Request Signing ────────────────────────────────────────── + +def generate_headers( + app_id: str, app_secret: str, + access_token: str, body: dict | None = None, + is_form: bool = False, +) -> dict: + """Generate signed request headers per Bilibili Open Platform spec.""" + md5_str = json.dumps(body) if body else "" + x_bili_content_md5 = hashlib.md5(md5_str.encode()).hexdigest() + + headers = { + "Accept": "application/json", + "Content-Type": "multipart/form-data" if is_form else "application/json", + "x-bili-content-md5": x_bili_content_md5, + "x-bili-timestamp": str(int(time.time())), + "x-bili-signature-method": "HMAC-SHA256", + "x-bili-signature-nonce": str(uuid.uuid4()), + "x-bili-accesskeyid": app_id, + "x-bili-signature-version": "2.0", + "access-token": access_token, + "Authorization": "", + } + + # Sort x-bili-* headers, join with \n, HMAC-SHA256 sign + header_str = "\n".join( + f"{k}:{headers[k]}" + for k in sorted(k for k in headers if k.startswith("x-bili-")) + ) + signature = hmac.new( + app_secret.encode(), header_str.encode(), hashlib.sha256, + ).hexdigest() + headers["Authorization"] = signature + + return headers + + +# ── API Operations ─────────────────────────────────────────────────────── + +def video_init(access_token: str, filename: str) -> str: + """Initialize video upload, return upload_token.""" + app_id, app_secret = get_app_credentials() + body = {"name": filename, "utype": "0"} + headers = generate_headers(app_id, app_secret, access_token, body=body) + resp = requests.post( + f"{MEMBER_BASE}/archive/video/init", + headers=headers, json=body, timeout=30, + ) + data = resp.json() + if data.get("code") != 0: + err_exit(f"UPLOAD_FAILED: video init: {data.get('message', data)}") + return data["data"]["upload_token"] + + +def upload_chunks(access_token: str, video_path: str, upload_token: str) -> list[dict]: + """Upload video in chunks, return list of {part_number, etag}.""" + app_id, app_secret = get_app_credentials() + headers = generate_headers(app_id, app_secret, access_token) + # Remove Content-Type for binary upload + upload_headers = {k: v for k, v in headers.items() if k != "Content-Type"} + + parts = [] + with open(video_path, "rb") as f: + part_num = 0 + while True: + chunk = f.read(CHUNK_SIZE) + if not chunk: + break + part_num += 1 + resp = requests.post( + f"{UPLOAD_BASE}/part/upload", + headers=upload_headers, + params={"upload_token": upload_token, "part_number": part_num}, + data=chunk, + timeout=120, + ) + data = resp.json() + if data.get("code") != 0: + err_exit(f"UPLOAD_FAILED: chunk {part_num}: {data.get('message', data)}") + etag = data["data"]["etag"] + parts.append({"part_number": part_num, "etag": etag}) + sys.stderr.write(f"[bilibili-publish] uploaded chunk {part_num}\n") + return parts + + +def video_complete(access_token: str, upload_token: str) -> None: + """Complete chunked upload (server-side merge).""" + app_id, app_secret = get_app_credentials() + headers = generate_headers(app_id, app_secret, access_token) + resp = requests.post( + f"{MEMBER_BASE}/archive/video/complete", + headers=headers, + params={"upload_token": upload_token}, + timeout=30, + ) + data = resp.json() + if data.get("code") != 0: + err_exit(f"UPLOAD_FAILED: video complete: {data.get('message', data)}") + + +def cover_upload(access_token: str, cover_path: str) -> str: + """Upload cover image, return cover URL.""" + app_id, app_secret = get_app_credentials() + headers = generate_headers(app_id, app_secret, access_token, is_form=True) + with open(cover_path, "rb") as f: + files = {"file": (os.path.basename(cover_path), f, "image/jpeg")} + resp = requests.post( + f"{MEMBER_BASE}/archive/cover/upload", + headers=headers, files=files, timeout=30, + ) + data = resp.json() + if data.get("code") != 0: + err_exit(f"UPLOAD_FAILED: cover: {data.get('message', data)}") + return data["data"]["url"] + + +def archive_add( + access_token: str, upload_token: str, + title: str, desc: str, cover_url: str, + tid: int, tags: str, copyright_type: int, +) -> dict: + """Submit video archive.""" + app_id, app_secret = get_app_credentials() + body = { + "title": title, + "desc": desc or "", + "cover": cover_url, + "tid": tid, + "tag": tags, + "copyright": copyright_type, + } + headers = generate_headers(app_id, app_secret, access_token, body=body) + resp = requests.post( + f"{MEMBER_BASE}/archive/add-by-utoken", + headers=headers, + params={"upload_token": upload_token}, + json=body, + timeout=30, + ) + data = resp.json() + if data.get("code") != 0: + err_exit(f"SUBMIT_FAILED: {data.get('message', data)}") + + resource_id = data["data"].get("resource_id", "") + # resource_id format: "avid:bvid" or just "avid" + bvid = "" + if ":" in str(resource_id): + _, bvid = str(resource_id).split(":", 1) + return {"ok": True, "resource_id": resource_id, "bvid": bvid, "url": f"https://www.bilibili.com/video/{bvid}" if bvid else ""} + + +# ── Main ────────────────────────────────────────────────────────────────── + +def main() -> None: + parser = argparse.ArgumentParser(description="Publish video to Bilibili via Open Platform API") + parser.add_argument("--title", required=True, help="Video title (max 80 chars)") + parser.add_argument("--video", required=True, help="Video file path") + parser.add_argument("--cover", help="Cover image path") + parser.add_argument("--desc", default="", help="Video description") + parser.add_argument("--tid", type=int, default=122, help="Partition ID (default: 122=野生技术协会)") + parser.add_argument("--tags", required=True, help="Comma-separated tags") + parser.add_argument("--copyright", type=int, default=1, choices=[1, 2], help="1=self-made, 2=repost") + parser.add_argument("--exchange-token", help="Exchange OAuth authorization code for access token") + args = parser.parse_args() + + # OAuth code exchange mode + if args.exchange_token: + exchange_token(args.exchange_token) + return + + # Validate video file + video_path = os.path.abspath(args.video) + if not os.path.isfile(video_path): + err_exit(f"UPLOAD_FAILED: video not found: {video_path}") + + if len(args.title) > 80: + err_exit("TITLE_TOO_LONG: title exceeds 80 characters") + + # Get access token (auto-refresh if needed) + access_token = get_access_token() + + # Step 1: Initialize upload + filename = os.path.basename(video_path) + sys.stderr.write("[bilibili-publish] initializing upload...\n") + upload_token = video_init(access_token, filename) + + # Step 2: Upload chunks + sys.stderr.write("[bilibili-publish] uploading chunks...\n") + upload_chunks(access_token, video_path, upload_token) + + # Step 3: Complete upload + sys.stderr.write("[bilibili-publish] completing upload...\n") + video_complete(access_token, upload_token) + + # Step 4: Upload cover (optional) + cover_url = "" + if args.cover and os.path.isfile(args.cover): + sys.stderr.write("[bilibili-publish] uploading cover...\n") + cover_url = cover_upload(access_token, args.cover) + + # Step 5: Submit archive + sys.stderr.write("[bilibili-publish] submitting archive...\n") + result = archive_add( + access_token, upload_token, + args.title, args.desc, cover_url, + args.tid, args.tags, args.copyright, + ) + + output(result) + + +if __name__ == "__main__": + main() diff --git a/addons/officials/crew/selfmedia-operator/skills/content-calibrator/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/SKILL.md new file mode 100644 index 00000000..437191f2 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/SKILL.md @@ -0,0 +1,481 @@ +--- +name: content-calibrator +description: 内容校准预测循环——打分 → 盲预测 → T+3d复盘 → 进化 rubric。按平台独立迭代,每个平台拥有自己的 rubric、校准池、预测日志、受众画像。本技能负责打分(blind sub-agent + score-only.sh + 阈值门 + 流程 1A)与校准闭环;发布记录与数据采集由 published-track 统一管理。 +metadata: + openclaw: + emoji: 🎯 + requires: + bins: + - bash + - sqlite3 + - node +--- + +# Content Calibrator — 内容校准预测循环 + +> 方法论源自 cheat-on-content,适配 openclaw + selfmedia-operator 工作流。 +> **三条不可妥协原则**: +> 1. **盲预测**:预测必须在看到实际数据之前写完,写完即 immutable +> 2. **升级 = 全量重打**:rubric 升级时校准池所有样本必须重打分 +> 3. **rubric 是工作台不是博物馆**:被推翻/吸收的观察删掉,git history 是档案 + +--- + +## 核心设计:按平台独立迭代 + +**一套程序逻辑,N 套校准实例。** 每个平台拥有完全独立的: + +| 组件 | 说明 | 为什么必须分开 | +|------|------|---------------| +| rubric | 评分维度 + 权重 | 不同内容形态的预测维度不同;即使维度相同,权重会分化 | +| calibration pool | 校准样本池 | 不同平台的 baseline 量级不同(公众号 1w 阅读 ≠ 小红书 1w 浏览) | +| predictions | 预测日志 | 同一篇内容发到两个平台,预测是两条独立 immutable 日志 | +| baseline | 基线播放/阅读量 | 平台间量级差异巨大 | +| bucket | 流量档位边界 | 绝对桶完全不同,比率桶的 baseline 也不同 | +| benchmark | 对标账号 | 公众号对标和小红书对标的 pattern 完全不同 | +| audience | 受众画像 | 两个平台受众画像差异很大 | + +--- + +## 核心闭环 + +``` +📊 打分 → 🎯 盲预测 → 🚀 发布 → 📈 T+3d复盘 → 🧬 进化 rubric +``` + +每个平台独立走这个闭环。公众号的 rubric 升级不影响小红书的 rubric。 + +--- + +## 与 published-track 的集成 + +发布流程为 **打分(1A) → 发布 → 记录(1B)**。**打分(1A)由本技能负责,发布记录(1B)由 published-track 负责。** + +### 流程 1A·打分评估(发布前自检) + +发布前对稿件做盲打分 + 阈值门,**避免主 agent 自创自评**。 + +1. **主 agent `sessions_spawn` 一个 blind sub-agent**,只喂 `script_path`(稿件/视频定稿)+ `calibration//rubric_notes.md`。sub-agent 硬禁读 `.cheat-state.json`/`predictions/`/`rubric-memo.md`/`audience.md`/对话历史,输出严格 JSON 7 维分(ER/HP/SR/QL/NA/AB/PV,各 0-5)+ per-dim confidence。 +2. 主 agent 拿分调 `score-only.sh` 校验 + 算 composite + 判阈值门: + ```bash + ./skills/content-calibrator/scripts/score-only.sh \ + --platform wx_mp --content-path "output_articles/xxx/article.md" \ + --cal-er 3 --cal-hp 4 --cal-sr 3 --cal-ql 4 --cal-na 3 --cal-ab 4 --cal-pv 2 + ``` + 返回 JSON 含 `passed` 与 `failing_dims`。阈值取自 `calibration//.cheat-state.json` 的 `score_threshold`(默认 0=不拦截),**每维需 > 阈值**才算通过。 +3. **阈值门**:`passed=false` → 主 agent 据 `failing_dims` 改稿 → 重新 spawn blind sub-agent 打分 → 再判门。**最多 2 轮**,仍不达标 → 暂停发布、上报用户裁定。 +4. `passed=true` → 放行,进入发布技能。 +5. **平台未启用 calibration**(`calibration//` 不存在)→ 跳过 1A,直接发布。 + +> **视频内容**:打分对象是**脚本定稿**(storyboard/口播稿),不是成片。视频技能流程 = 打分(定稿) → 制作 → 发布 → 记录。成片后不再打分。 + +### 流程 1B·发布记录(由 published-track 承接) + +打分通过并发布成功后,由 `published-track/scripts/record.sh` 落库(提供 `--cal-*` 分数 → `cal_enabled=1` + 算 composite;不提供 → `cal_enabled=0`)。详见 `published-track/SKILL.md`。 + +### 平台打分开关 + 阈值 + +`content-calibrator/scripts/cal-toggle.sh`(`--enable/--disable/--set-threshold N/--threshold/--list`)。 + +### 数据采集由 published-track 统一管理 + +**content-calibrator 不再直接抓取平台数据。** 数据采集流程: + +1. **一键获取**:使用 `published-track/scripts/fetch-and-update-metrics.sh`(封装 login-manager 探活 → API 抓取 → DB 写入) +2. **复盘时**:直接从 published-track DB 读取数据,不另行抓取 +3. **深度数据**(完播率、转粉率、评论内容等):仍需 browser tool + CDP 拦截,但由 published-track 的心跳任务负责采集和更新 + +### wx-mp-hunter 数据获取说明 + +⚠️ **wx-mp-hunter 无法获取微信公众号文章的阅读数、点赞数等互动数据。** 微信公众号的互动数据需要登录微信公众号后台查看,wx-mp-hunter 只能获取文章标题、正文和链接。 + +因此: +- 微信公众号的复盘数据只能从 published-track DB 中读取(由心跳巡检手动更新) +- 如需精确数据,用户需手动提供或通过微信公众号后台截图 + +--- + +## 路由表(触发词 → 操作) + +| 用户说 | 操作 | 前置条件 | +|--------|------|----------| +| "初始化校准 [--platform xxx]" | Init | 首次使用 | +| "打分这篇 [path] --platform xxx" | Score | 该平台 rubric_notes.md 存在 | +| "预测这篇 [path] --platform xxx" / "启动预测" | Predict | 已 init + 有最终稿 | +| "复盘 [path] --platform xxx" / "T+3d 数据来了" | Retro | 有预测 + 已发布 + 过时间窗口 | +| "升级公式 --platform xxx" / "bump rubric" | Bump | 校准池 ≥ MIN_SAMPLES | +| "导入对标 --platform xxx" / "learn from" | LearnFrom | 有 viral-chaser 报告或用户提供对标数据 | +| "校准状态 [--platform xxx]" / "calibration status" | Status | 任意时刻 | +| "加维度 XX --platform xxx" | 维度变更 | **必须用户确认** | +| "改权重 XX --platform xxx" | 权重变更 | **必须用户确认** | + +### 平台启用控制 + +**是否启用某个平台的 calibration,必须由用户决定。** Agent 不得自动启用。 + +- 启用:`./skills/content-calibrator/scripts/cal-toggle.sh --platform --enable` +- 停用:`./skills/content-calibrator/scripts/cal-toggle.sh --platform --disable` +- 查看状态:`./skills/content-calibrator/scripts/cal-toggle.sh --list` + +Agent 在复盘或发布时,发现对应平台未启用 calibration,**不得自动启用**,应告知用户"该平台未启用 content-calibrator,如需启用请确认"。 + +`--platform` 为必填参数(Init 除外,Init 时交互询问或 Agent 自主判断)。支持的平台 ID: + +| 平台 ID | 平台 | 内容形态 | +|---------|------|---------| +| `wx_mp` | 微信公众号 | 长文 | +| `wx_channel` | 微信视频号 | 短视频 | +| `xhs` | 小红书 | 图文/视频笔记 | +| `zhihu` | 知乎 | 文章/回答 | +| `bilibili` | B站 | 视频 | +| `douyin` | 抖音 | 短视频 | +| `kuaishou` | 快手 | 短视频 | +| `toutiao` | 今日头条 | 文章 | +| `youtube` | YouTube | 视频 | + +--- + +## 文件结构 + +``` +/ +├── calibration/ # 校准系统根目录 +│ ├── wx_mp/ # 公众号独立体系 +│ │ ├── rubric_notes.md # 评分公式(blind sub-agent 可读) +│ │ ├── rubric-memo.md # 观察记录(含实绩数据,blind 不可读) +│ │ ├── .cheat-state.json # 该平台状态文件 +│ │ ├── predictions/ # immutable 预测日志 +│ │ │ └── YYYY-MM-DD__.md +│ │ ├── benchmark.md # 对标账号信息 +│ │ └── audience.md # 受众画像 +│ ├── wx_channel/ # 视频号独立体系 +│ │ ├── rubric_notes.md +│ │ ├── rubric-memo.md +│ │ ├── .cheat-state.json +│ │ ├── predictions/ +│ │ ├── benchmark.md +│ │ └── audience.md +│ ├── xhs/ # 小红书独立体系 +│ │ ├── rubric_notes.md +│ │ ├── rubric-memo.md +│ │ ├── .cheat-state.json +│ │ ├── predictions/ +│ │ ├── benchmark.md +│ │ └── audience.md +│ └── ... # 更多平台 +``` + +--- + +## Init — 初始化 + +为指定平台创建 `calibration//` 目录和初始文件。 + +**两种触发方式**: +- **用户主动**:用户说"初始化校准"或"我要做 XX 平台" → 交互式问答 +- **Agent 不得自主初始化**:必须用户明确要求 + +### 用户主动触发流程 + +1. 询问或从 `--platform` 参数获取平台 ID +2. 创建目录结构 `calibration//` +3. 写入 `rubric_notes.md`(根据平台选择默认 rubric) +4. 写入 `.cheat-state.json`(cold-start 模式) +5. 询问用户 5 个问题:内容形态、典型篇幅、发布频率、对标账号(可选)、该平台 baseline +6. 如有对标账号 → 触发 LearnFrom + +### 默认 Rubric 按平台分发 + +| 平台 | 默认 rubric | 说明 | +|------|------------|------| +| wx_mp | 长文 rubric(ER/HP/SR×1.5 + QL/NA/AB/PV×1.0) | 公众号核心驱动力:情感 + 钩子 + 社会议题 | +| wx_channel | 视频 rubric(ER/HP/SR×1.5 + QL/NA/AB/PV×1.0) | 视频号核心驱动力:情感 + 钩子 + 社会议题 | +| xhs | 图文笔记 rubric(ER/HP/SR×1.5 + QL/NA/AB/PV×1.0) | 起步同维度,bump 后会分化 | +| bilibili/douyin/kuaishou | 视频 rubric(ER/HP/SR×1.5 + QL/NA/AB/PV×1.0) | 起步同维度,bump 后会分化 | +| 其他 | 同上 | 起步同维度,bump 后会分化 | + +--- + +## Score — 打分 + +给单篇稿子打 rubric 分。**脚本不做 LLM 打分**;打分由主 agent `sessions_spawn` 的 blind sub-agent 完成,在发布前作为自检门(流程见上方"流程 1A·打分评估")。 + +**blind sub-agent 隔离规则**(主对话已看过用户对话/实绩/复盘历史,inline 打分会被污染,故必须 delegate): + +- **白名单只读**:稿件(`script.md`/`article.md`/`post.md`)+ `calibration//rubric_notes.md` +- **硬禁读**:`rubric-memo.md`、`.cheat-state.json`、`predictions/`、`audience.md`、`benchmark.md`、对话历史 +- **输出**:严格 JSON 7 维分(ER/HP/SR/QL/NA/AB/PV,各 0-5)+ per-dim confidence +- Bump Phase 2 校准池重打分**强制** blind sub-agent,不接受任何 fallback +- Predict Phase 2.5 做 disagreement detection:blind 与主 Claude 自估 |delta| ≥ 2 → 弹用户裁定 + +### 当前默认 rubric(所有平台起步版 v0) + +7 个维度,每维 0-5 整数分: + +| 维度 | 代号 | 含义 | 权重 | +|------|------|------|------| +| 情感共鸣 | ER | 读者能否产生"说的就是我"的代入感 | ×1.5 | +| 钩子强度 | HP | 标题/开头是否锁定注意力 | ×1.5 | +| 社会议题共振 | SR | 是否触及社会讨论 | ×1.5 | +| 金句密度 | QL | 是否有独立可传播的表达 | ×1.0 | +| 叙事性 | NA | 是否有清晰的故事弧线 | ×1.0 | +| 受众广度 | AB | 话题的普适程度 | ×1.0 | +| 实用价值 | PV | 读者能否获得可操作的信息 | ×1.0 | + +**composite = (ER×1.5 + HP×1.5 + SR×1.5 + QL + NA + AB + PV) / 8.5 × 2.0** + +--- + +## Predict — 盲预测 + +在看到任何实际数据之前写 immutable 预测日志。 + +### 流程 + +1. **Blind check**:确认用户未看过后续数据 +2. 读最终稿 + `calibration//rubric_notes.md` + state +3. **Blind sub-agent 盲打分**(只看稿子 + rubric,不看对话历史) +4. 锚点对比(从 `calibration//predictions/` 找历史相近 composite 的样本) +5. 给 bucket(流量档位)+ 概率分布 + 中枢 +6. 写反事实场景 + 关键校准假设 +7. **用户 review**:展示完整草拟版,等 "ok" 或挑刺 +8. 落盘到 `calibration//predictions/`,预测段 immutable + +### Cold-start 简化 + +前 5 篇不要求完整 bucket 数字,只给 7 维分 + 一句话 bet。第 5 篇复盘后解锁完整预测。 + +--- + +## Retro — 复盘 + +T+N 天后从 published-track DB 读实际数据 → 对比预测 → 提炼观察。 + +### 两个入口 + +#### 入口 1:凌晨 HEARTBEAT 自动复盘 + +心跳巡检时,检查每个已启用 calibration 的平台: +- 从 published-track DB 读取该平台所有 `cal_enabled=1` 且有预测但未复盘的记录 +- 检查是否积累了 **5 个新数据点**(有实际互动数据但尚未复盘) +- 如有 ≥5 个 → 自动执行复盘流程 + +#### 入口 2:用户导入对标 + +用户主动提供对标账号/爆款内容数据,触发 LearnFrom 流程。这是**校准 rubric 本身**的入口——通过分析对标内容,提炼该平台高流量内容的 pattern,调整 rubric 维度和权重。 + +> **复盘的本质**:复盘是"拿实际数据验证预测,提炼观察,可能触发 rubric 升级"。导入对标是"从外部信号校准 rubric 的初始假设"。两者互补:复盘是内源校准,对标是外源校准。 + +### 数据来源(全部从 published-track DB) + +复盘时**只从 published-track DB 读取数据**,不另行抓取: + +```bash +# 读取某平台某篇内容的互动指标 +./skills/published-track/scripts/query.sh --platform wx_mp --limit 10 + +# 或直接 SQL +sqlite3 db/published_track.db "SELECT * FROM pub_wx_mp WHERE source_folder='output_articles/xxx'" +``` + +### 阈值推荐(复盘副产物) + +复盘积累数据后,Agent 可评估"发布前自检阈值门"的 `score_threshold` 是否合理:观察各维度分与实际互动的相关性,若某维度低分内容普遍表现差,可建议提高该平台阈值。**Agent 不得自动改阈值**,需向用户给出建议值与依据,经用户确认后执行: + +```bash +./skills/content-calibrator/scripts/cal-toggle.sh --platform --set-threshold +``` + +起步期阈值默认 0(不拦截),待累积足够复盘样本后再收紧。 + +**各平台数据获取能力**(由 `published-track/scripts/fetch-and-update-metrics.sh` 统一调度): + +| 平台 | 方案 | 心跳可自动更新 | 需手动补充 | 说明 | +|------|------|--------------|-----------|------| +| 小红书 | 脚本 | ✅ likes/favorites/comments/shares | 完播率/转粉率 | xhshow 签名 + cookie | +| B站 | 脚本 | ✅ plays/likes/coins/favorites/danmaku/comments | 完播率 | 公开 API,无需 cookie | +| 抖音 | 脚本 | ⚠️ plays/likes/comments/shares/favorites | 完播率 | a_bogus 签名 + cookie | +| 快手 | 脚本 | ⚠️ plays/likes/comments | — | GraphQL + cookie | +| 知乎 | 浏览器 | ⚠️ views/upvotes/comments/favorites | — | browser snapshot | +| 今日头条 | 浏览器 | ⚠️ impressions/reads/comments/likes | — | browser snapshot | +| 掘金 | 浏览器 | ✅ views/likes/comments/favorites | — | browser snapshot(无需 cookie) | +| Twitter/X | 浏览器 | ⚠️ views/likes/retweets/replies/bookmarks | — | twitter-interact 技能 | +| YouTube | 浏览器 | ✅ views/likes/comments/shares | — | browser snapshot(无需 cookie) | +| 微信公众号 | 跳过 | ❌ 无法自动获取 | reads/likes/shares 等 | 需用户手动提供 | +| 微信视频号 | 跳过 | ❌ 无法自动获取 | plays/likes/comments/shares/favorites | 需用户手动提供 | + +### 流程 + +1. 校验时间窗口(默认 T+3d) +2. 从 published-track DB 读互动数据 +3. 写实绩段 + top 评论关键词聚类(如有评论数据) +4. 验证/推翻预测的各假设 +5. 提炼新观察 → 写入 `calibration//rubric-memo.md` +6. 检测是否触发 bump(≥3 次同向偏差) + +--- + +## Bump — Rubric 升级 + +系统性偏差信号 → 校准池全量重打 → 排序一致性校验 → 落地新公式。 + +**只影响当前平台的 rubric**,其他平台不受影响。 + +### 流程 + +1. 前置门槛检查(校准池样本数 + 观察强度) +2. 写出新公式完整方程 +3. 校准池全量重打分(blind sub-agent 隔离) +4. 计算排序一致性(新公式排序 vs 实际排序,阈值 4/5) +5. 落地 + cleanup pass(删被推翻/吸收的观察) +6. 更新所有校准样本的 Re-scored 标记 +7. 更新 `calibration//rubric_notes.md` 版本速查 + +--- + +## 维度与权重变更规则 + +**维度和权重可以被修改,但必须满足以下条件之一**: +1. **用户主动要求** — "给公众号加个 XX 维度" / "把小红书的 SR 权重调到 2.0" +2. **Agent 提议 + 用户确认** — Agent 在 Bump 流程中检测到系统性偏差后提议变更,**必须等待用户明确同意才生效** + +变更流程: +- 变更维度(增/删/替换)→ 走 Bump 全量重打 + 排序一致性校验 +- 变更权重 → 走 Bump 流程 +- 变更被拒绝 → rubric 不动,观察记入 `rubric-memo.md` + +--- + +## LearnFrom — 导入对标 + +从对标账号/爆款内容中提取 pattern,作为该平台 rubric 初始校准信号。 + +### 数据来源 + +1. **viral-chaser 追爆报告**:已下载的爆款视频分析 → 提取结构 pattern +2. **用户提供的数据**:手动粘贴对标账号数据 +3. **published-track DB 中的历史数据**:该平台已发布内容的互动数据 + +> ⚠️ wx-mp-hunter **无法获取**公众号互动数据,不能作为对标数据来源。 + +### 流程 + +1. 确认对标来源(viral-chaser 报告 / 用户提供数据 / 历史数据) +2. 分析 pattern:哪些维度在高流量内容中一致偏高/偏低 +3. 派生 rubric 信号(调整权重/维度) +4. 写入 `calibration//benchmark.md` + 更新 `rubric-memo.md` + +--- + +## Status — 校准状态看板 + +显示指定平台(或所有平台)的校准循环状态: + +``` +📊 Content Calibrator 状态 + +平台: wx_mp(微信公众号) +模式: calibration(已过 cold-start) +Rubric: v0(长文) +校准池: 8 篇(5 篇有实绩) +上次预测: 2026-06-08 +待复盘: 2 篇 +Buffer: 0 + +最近 5 篇偏差: + ✅ AI团队运营 预测 30-100w 实际 45w +13% + ❌ 一人公司 预测 30-100w 实际 8w -73% + ✅ 搞钱思维 预测 5-30w 实际 22w +10% + ... + +系统性偏差: 无(连续同向 < 3) + +--- + +平台: xhs(小红书) +模式: cold-start +Rubric: v0(图文笔记) +校准池: 0 篇 +(尚未开始校准循环) +``` + +--- + +## 脚本 + +### 发布记录(合并入口 record.sh) + +Agent 发布后调用 `record.sh`,行为:提供任意 `--cal-*` 分数 → 置 `cal_enabled=1` + 算 composite + 记 `cal_rubric_version`;不提供 → `cal_enabled=0`。`score-and-record.sh` 已合并为 `record.sh` 的薄 wrapper。 + +```bash +# 打分通过的平台 — 把 blind sub-agent 打的分一并传入 +./skills/published-track/scripts/record.sh \ + --platform wx_mp \ + --title "AI 时代的一人公司" \ + --content-type article \ + --source-folder "output_articles/ai-one-person-company" \ + --publish-url "https://mp.weixin.qq.com/s/xxx" \ + --cal-er 3 --cal-hp 4 --cal-sr 4 --cal-ql 3 --cal-na 2 --cal-ab 4 --cal-pv 3 + +# 未启用 calibration 或补发 — 不传 --cal-* 即可 +./skills/published-track/scripts/record.sh \ + --platform xhs \ + --title "标题" \ + --content-type post \ + --source-folder "output_articles/xxx" \ + --publish-url "https://www.xiaohongshu.com/xxx" +``` + +### 打分结果校验(不写入数据库) + +Agent 按 rubric 打完 7 维分后,可用 `score-only.sh` 校验分数合法性、计算 composite 并输出结构化 JSON,不写入 DB。此脚本不做 LLM 打分,仅校验并格式化 Agent 已打好的分数。 + +```bash +./skills/content-calibrator/scripts/score-only.sh \ + --platform wx_mp \ + --content-path "output_articles/xxx/article.md" \ + --cal-er 3 --cal-hp 4 --cal-sr 4 --cal-ql 3 --cal-na 2 --cal-ab 4 --cal-pv 3 +``` + +### 平台打分开关管理 + +```bash +# 查看所有平台 +./skills/content-calibrator/scripts/cal-toggle.sh --list + +# 启用/停用 +./skills/content-calibrator/scripts/cal-toggle.sh --platform wx_mp --enable +./skills/content-calibrator/scripts/cal-toggle.sh --platform wx_mp --disable +``` + +### 初始化平台 + +```bash +./skills/content-calibrator/scripts/init.sh --platform +``` + +创建 `calibration//` 目录和初始文件。幂等——已存在则跳过。 + +### 查询 published-track 数据 + +```bash +./skills/content-calibrator/scripts/query-metrics.sh --platform --source-folder +``` + +从 published-track DB 查询某篇内容的互动指标。 + +### 构建校准池 + +```bash +./skills/content-calibrator/scripts/build-calibration-pool.sh --platform +``` + +从 published-track DB 构建指定平台的校准池。 + +### 导入追爆报告 + +```bash +./skills/content-calibrator/scripts/import-viral-chaser.sh --platform +``` + +将 viral-chaser 追爆报告导入为指定平台的对标信号。 diff --git a/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/build-calibration-pool.sh b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/build-calibration-pool.sh new file mode 100755 index 00000000..00d8a427 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/build-calibration-pool.sh @@ -0,0 +1,77 @@ +#!/usr/bin/env bash +# build-calibration-pool.sh — 从 published-track DB 构建指定平台的校准池 +# 用法: build-calibration-pool.sh --platform +# 输出该平台有互动数据的发布记录,供复盘和 bump 使用 +set -euo pipefail + +WORKSPACE="$( cd -- "$( dirname -- "${BASH_SOURCE[0]}" )/../../.." &> /dev/null && pwd )" +DB="$WORKSPACE/db/published_track.db" +CAL_ROOT="$WORKSPACE/calibration" + +PLATFORM="" + +while [[ $# -gt 0 ]]; do + case "$1" in + --platform) PLATFORM="$2"; shift 2 ;; + *) echo "未知参数: $1"; exit 1 ;; + esac +done + +if [[ -z "$PLATFORM" ]]; then + echo "用法: build-calibration-pool.sh --platform " + echo " platform: wx_mp | xhs | zhihu | bilibili | douyin | kuaishou | toutiao | youtube" + exit 1 +fi + +CAL_DIR="$CAL_ROOT/$PLATFORM" +if [[ ! -d "$CAL_DIR" ]]; then + echo "❌ 平台 $PLATFORM 的校准目录不存在: $CAL_DIR" + echo " 先运行 init.sh --platform $PLATFORM" + exit 1 +fi + +if [[ ! -f "$DB" ]]; then + echo "❌ published-track DB 不存在: $DB" + exit 1 +fi + +TABLE="pub_${PLATFORM}" + +# 检查表是否存在 +table_exists=$(sqlite3 "$DB" "SELECT count(*) FROM sqlite_master WHERE type='table' AND name='$TABLE';") +if [[ "$table_exists" -eq 0 ]]; then + echo "❌ 平台表不存在: $TABLE" + exit 1 +fi + +# 平台 → 主指标字段映射 +declare -A METRIC_FIELD +METRIC_FIELD[wx_mp]="reads" +METRIC_FIELD[xhs]="views" +METRIC_FIELD[zhihu]="views" +METRIC_FIELD[bilibili]="plays" +METRIC_FIELD[douyin]="plays" +METRIC_FIELD[kuaishou]="plays" +METRIC_FIELD[toutiao]="reads" +METRIC_FIELD[youtube]="views" + +METRIC="${METRIC_FIELD[$PLATFORM]:-views}" + +echo "📊 从 published_track.$TABLE 构建校准池(主指标: $METRIC)..." +echo "" + +# 查询该平台有互动数据的记录 +sqlite3 -separator "|" "$DB" " +SELECT title, source_folder, publish_date, + COALESCE($METRIC, 0) as metric +FROM $TABLE WHERE $METRIC > 0 +ORDER BY publish_date DESC; +" 2>/dev/null | while IFS='|' read -r title folder date metric; do + echo "$title | $folder | $date | $metric" +done + +count=$(sqlite3 "$DB" "SELECT count(*) FROM $TABLE WHERE $METRIC > 0;" 2>/dev/null || echo "0") + +echo "" +echo "---" +echo "平台: $PLATFORM | 总计: $count 条有互动数据的记录" diff --git a/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/cal-toggle.sh b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/cal-toggle.sh new file mode 100755 index 00000000..27d388dc --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/cal-toggle.sh @@ -0,0 +1,124 @@ +#!/usr/bin/env bash +# cal-toggle.sh — 管理平台 content-calibrator 打分开关与阈值 +# +# 用法: +# cal-toggle.sh --list # 查看所有平台打分开关状态 +# cal-toggle.sh --platform --enable # 启用某平台打分 +# cal-toggle.sh --platform --disable # 停用某平台打分 +# cal-toggle.sh --platform --status # 查看某平台打分状态 +# cal-toggle.sh --platform --threshold # 查看某平台打分阈值 +# cal-toggle.sh --platform --set-threshold # 设置阈值(每维 0-5,需 >N 才放行发布;0=不拦截) +set -euo pipefail + +ROOT="$(cd "$(dirname "$0")/../../.." && pwd)" +CAL_ROOT="$ROOT/calibration" +DB="$ROOT/db/published_track.db" + +ACTION="" PLATFORM="" THRESHOLD_VAL="" + +while [[ $# -gt 0 ]]; do + case "$1" in + --platform) PLATFORM="$2"; shift 2 ;; + --enable) ACTION=enable; shift ;; + --disable) ACTION=disable; shift ;; + --status) ACTION=status; shift ;; + --list) ACTION=list; shift ;; + --threshold) ACTION=threshold; shift ;; + --set-threshold) ACTION=set_threshold; THRESHOLD_VAL="$2"; shift 2 ;; + *) echo "{\"ok\":false,\"error\":\"unknown arg: $1\"}"; exit 1 ;; + esac +done + +# 支持的平台 +VALID_PLATFORMS="wx_mp wx_channel xhs zhihu bilibili douyin kuaishou toutiao youtube juejin twitter facebook instagram tiktok pinterest threads" + +if [ "$ACTION" = "list" ]; then + echo "📊 Content-Calibrator 平台打分开关" + echo "" + for p in $VALID_PLATFORMS; do + CAL_DIR="$CAL_ROOT/$p" + if [ -d "$CAL_DIR" ] && [ -f "$CAL_DIR/.cheat-state.json" ]; then + thr=$(python3 -c "import json; print(json.load(open('$CAL_DIR/.cheat-state.json')).get('score_threshold',0))" 2>/dev/null || echo 0) + echo " ✅ $p — 已启用(阈值 >$thr 放行)" + else + echo " ⬜ $p — 未启用" + fi + done + exit 0 +fi + +if [ -z "$PLATFORM" ]; then + echo '{"ok":false,"error":"--platform is required (or use --list)"}' + exit 1 +fi + +if ! echo "$VALID_PLATFORMS" | grep -qw "$PLATFORM"; then + echo "{\"ok\":false,\"error\":\"unsupported platform: $PLATFORM\"}" + exit 1 +fi + +CAL_DIR="$CAL_ROOT/$PLATFORM" + +case "$ACTION" in + status) + if [ -d "$CAL_DIR" ] && [ -f "$CAL_DIR/.cheat-state.json" ]; then + echo "{\"ok\":true,\"platform\":\"$PLATFORM\",\"cal_enabled\":true}" + else + echo "{\"ok\":true,\"platform\":\"$PLATFORM\",\"cal_enabled\":false}" + fi + ;; + threshold) + if [ ! -f "$CAL_DIR/.cheat-state.json" ]; then + echo "{\"ok\":false,\"error\":\"platform $PLATFORM not initialized (no .cheat-state.json)\"}"; exit 1 + fi + thr=$(python3 -c "import json; print(json.load(open('$CAL_DIR/.cheat-state.json')).get('score_threshold',0))") + echo "{\"ok\":true,\"platform\":\"$PLATFORM\",\"score_threshold\":$thr,\"meaning\":\"每维需 >$thr 才放行发布\"}" + ;; + set_threshold) + if [ ! -f "$CAL_DIR/.cheat-state.json" ]; then + echo "{\"ok\":false,\"error\":\"platform $PLATFORM not initialized (no .cheat-state.json)\"}"; exit 1 + fi + if [[ -z "$THRESHOLD_VAL" ]]; then + echo '{"ok":false,"error":"--set-threshold requires a value 0-4"}'; exit 1 + fi + if [[ "$THRESHOLD_VAL" -lt 0 || "$THRESHOLD_VAL" -gt 4 ]] 2>/dev/null; then + echo "{\"ok\":false,\"error\":\"threshold must be integer 0-4 (每维 0-5, 需 >threshold, 故 threshold 上限 4)\"}"; exit 1 + fi + python3 -c " +import json +f='$CAL_DIR/.cheat-state.json' +d=json.load(open(f)); d['score_threshold']=$THRESHOLD_VAL +json.dump(d,open(f,'w'),ensure_ascii=False,indent=2) +" + echo "{\"ok\":true,\"platform\":\"$PLATFORM\",\"score_threshold\":$THRESHOLD_VAL,\"meaning\":\"每维需 >$THRESHOLD_VAL 才放行发布\"}" + ;; + enable) + if [ -d "$CAL_DIR" ] && [ -f "$CAL_DIR/.cheat-state.json" ]; then + echo "{\"ok\":true,\"platform\":\"$PLATFORM\",\"action\":\"enable\",\"message\":\"already enabled\"}" + else + # 调 content-calibrator 的 init.sh + bash "$(dirname "$0")/../../content-calibrator/scripts/init.sh" --platform "$PLATFORM" + echo "{\"ok\":true,\"platform\":\"$PLATFORM\",\"action\":\"enable\",\"message\":\"calibration initialized\"}" + fi + ;; + disable) + if [ ! -d "$CAL_DIR" ]; then + echo "{\"ok\":true,\"platform\":\"$PLATFORM\",\"action\":\"disable\",\"message\":\"already disabled (dir not found)\"}" + else + echo "⚠️ 禁用 $PLATFORM 的 content-calibrator 将删除 calibration/$PLATFORM/ 目录" + echo " rubric、预测日志、对标数据等将全部删除" + echo " 确认请输入 YES: " + read -r CONFIRM + if [ "$CONFIRM" = "YES" ]; then + rm -rf "$CAL_DIR" + echo "{\"ok\":true,\"platform\":\"$PLATFORM\",\"action\":\"disable\",\"message\":\"calibration directory removed\"}" + else + echo "{\"ok\":false,\"platform\":\"$PLATFORM\",\"action\":\"disable\",\"message\":\"cancelled by user\"}" + fi + fi + ;; + *) + echo '{"ok":false,"error":"action required: --enable, --disable, --status, --threshold, --set-threshold, or --list"}' + exit 1 + ;; +esac diff --git a/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/import-viral-chaser.sh b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/import-viral-chaser.sh new file mode 100644 index 00000000..e4c99c51 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/import-viral-chaser.sh @@ -0,0 +1,90 @@ +#!/usr/bin/env bash +# import-viral-chaser.sh — 将 viral-chaser 追爆报告导入为指定平台的对标信号 +# 用法: import-viral-chaser.sh --platform +set -euo pipefail + +WORKSPACE="$( cd -- "$( dirname -- "${BASH_SOURCE[0]}" )/../../.." &> /dev/null && pwd )" +CAL_ROOT="$WORKSPACE/calibration" + +PLATFORM="" +REPORT_PATH="" + +while [[ $# -gt 0 ]]; do + case "$1" in + --platform) PLATFORM="$2"; shift 2 ;; + *) REPORT_PATH="$1"; shift ;; + esac +done + +if [[ -z "$PLATFORM" || -z "$REPORT_PATH" ]]; then + echo "用法: import-viral-chaser.sh --platform " + echo " platform_id: wx_mp | xhs | zhihu | bilibili | douyin | kuaishou | toutiao | youtube" + echo " report-path: viral-chaser 追爆报告.md 的路径" + echo "" + echo "示例:" + echo " import-viral-chaser.sh --platform douyin output_videos/douyin-7389abc/追爆报告.md" + exit 1 +fi + +CAL_DIR="$CAL_ROOT/$PLATFORM" +if [[ ! -d "$CAL_DIR" ]]; then + echo "❌ 平台 $PLATFORM 的校准目录不存在: $CAL_DIR" + echo " 先运行 init.sh --platform $PLATFORM" + exit 1 +fi + +if [[ ! -f "$REPORT_PATH" ]]; then + echo "❌ 报告文件不存在: $REPORT_PATH" + exit 1 +fi + +BENCHMARK_FILE="$CAL_DIR/benchmark.md" +if [[ ! -f "$BENCHMARK_FILE" ]]; then + echo "❌ benchmark.md 不存在,先运行 init.sh --platform $PLATFORM" + exit 1 +fi + +echo "🎯 导入追爆报告为对标信号 — 平台: $PLATFORM" +echo " 报告: $REPORT_PATH" + +# 提取报告关键信息 +REPORT_CONTENT=$(cat "$REPORT_PATH") + +# 提取平台 +REPORT_PLATFORM=$(echo "$REPORT_CONTENT" | grep -oP '(?<=平台[::]\s*)\S+' | head -1 || echo "unknown") +# 提取标题 +TITLE=$(echo "$REPORT_CONTENT" | grep -oP '(?<=标题[::]\s*).*' | head -1 || echo "unknown") +# 提取播放量 +PLAYS=$(echo "$REPORT_CONTENT" | grep -oP '(?<=播放[::]\s*)[\d.]+[wW万]?' | head -1 || echo "N/A") + +TIMESTAMP=$(date -Iseconds) + +# 追加到该平台的 benchmark.md +cat >> "$BENCHMARK_FILE" << EOF + +--- + +### 追爆对标 — $TITLE ($REPORT_PLATFORM) + +- **来源**: viral-chaser 追爆报告 +- **导入平台**: $PLATFORM +- **导入时间**: $TIMESTAMP +- **播放量**: $PLAYS +- **报告路径**: $REPORT_PATH + +**Pattern 提炼**(由 agent 从报告中分析): +- 结构 pattern: (待 agent 分析填充) +- 开头方式: (待分析) +- 转折技巧: (待分析) +- 金句模式: (待分析) +- 互动钩子: (待分析) + +**Rubric 信号**(对当前 rubric 维度的启示): +- (待 agent 从报告数据中提炼,如"高 ER + 高 HP → 高流量") + +EOF + +echo "" +echo "✅ 已追加到 calibration/$PLATFORM/benchmark.md" +echo "" +echo "下一步: 让 agent 分析追爆报告,填充 pattern 和 rubric 信号" diff --git a/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/init.sh b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/init.sh new file mode 100755 index 00000000..dd618da4 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/init.sh @@ -0,0 +1,96 @@ +#!/usr/bin/env bash +# content-calibrator init — 为指定平台创建校准系统目录和初始文件 +# 用法: init.sh --platform +# platform_id: wx_mp | xhs | zhihu | bilibili | douyin | kuaishou | toutiao | youtube +set -euo pipefail + +WORKSPACE="$( cd -- "$( dirname -- "${BASH_SOURCE[0]}" )/../../.." &> /dev/null && pwd )" +CAL_ROOT="$WORKSPACE/calibration" + +PLATFORM="" + +while [[ $# -gt 0 ]]; do + case "$1" in + --platform) PLATFORM="$2"; shift 2 ;; + *) echo "未知参数: $1"; exit 1 ;; + esac +done + +# 支持的平台列表 +VALID_PLATFORMS="wx_mp wx_channel xhs zhihu bilibili douyin kuaishou toutiao youtube" + +if [[ -z "$PLATFORM" ]]; then + echo "用法: init.sh --platform " + echo "" + echo "支持的平台:" + echo " wx_mp 微信公众号" + echo " wx_channel 微信视频号" + echo " xhs 小红书" + echo " zhihu 知乎" + echo " bilibili B站" + echo " douyin 抖音" + echo " kuaishou 快手" + echo " toutiao 今日头条" + echo " youtube YouTube" + exit 1 +fi + +if ! echo "$VALID_PLATFORMS" | grep -qw "$PLATFORM"; then + echo "❌ 不支持的平台: $PLATFORM" + echo " 支持的平台: $VALID_PLATFORMS" + exit 1 +fi + +CAL_DIR="$CAL_ROOT/$PLATFORM" + +echo "🔧 初始化 Content Calibrator — $PLATFORM" +echo " 工作区: $WORKSPACE" +echo " 校准目录: $CAL_DIR" +echo "" + +# 创建目录结构 +mkdir -p "$CAL_DIR/predictions" + +# 检查已有文件 +files=(rubric_notes.md rubric-memo.md .cheat-state.json benchmark.md audience.md) +existing=0 +for f in "${files[@]}"; do + if [[ -f "$CAL_DIR/$f" ]]; then + existing=$((existing + 1)) + fi +done + +if [[ $existing -eq ${#files[@]} ]]; then + echo "✅ 平台 $PLATFORM 的校准系统已初始化(所有文件均存在)" + echo "" + echo "当前状态:" + mode=$(python3 -c "import json; print(json.load(open('$CAL_DIR/.cheat-state.json'))['mode'])" 2>/dev/null || echo "unknown") + samples=$(python3 -c "import json; print(json.load(open('$CAL_DIR/.cheat-state.json'))['calibration_samples'])" 2>/dev/null || echo "0") + version=$(python3 -c "import json; print(json.load(open('$CAL_DIR/.cheat-state.json'))['rubric_version'])" 2>/dev/null || echo "unknown") + echo " 模式: $mode" + echo " Rubric: $version" + echo " 校准样本: $samples" + exit 0 +fi + +if [[ $existing -gt 0 ]]; then + echo "⚠️ 校准目录已存在部分文件($existing/${#files[@]}),跳过已存在文件。" +fi + +# 创建不存在的文件 +for f in "${files[@]}"; do + if [[ ! -f "$CAL_DIR/$f" ]]; then + echo " 创建 $f" + fi +done + +echo "" +echo "✅ 初始化完成 — 平台: $PLATFORM" +echo "" +echo "下一步:" +echo " 1. 对已有发布内容做首次复盘 → 积累校准样本" +echo " 2. 导入对标账号 → 获取初始 rubric 信号" +echo " 3. 对新稿子打分 → 开始校准循环" +echo "" +echo "其他平台初始化:" +echo " ./skills/content-calibrator/scripts/init.sh --platform <另一个平台>" diff --git a/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/query-metrics.sh b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/query-metrics.sh new file mode 100755 index 00000000..b83ac953 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/query-metrics.sh @@ -0,0 +1,50 @@ +#!/usr/bin/env bash +# query-metrics.sh — 从 published-track DB 查询某篇内容的互动指标 +# 用法: query-metrics.sh --platform --source-folder +# platform 对应 published-track 的平台 ID(wx_mp/xhs/zhihu/bilibili/douyin/kuaishou/toutiao/juejin/twitter/facebook/instagram/tiktok/youtube/pinterest/threads/wxwork_moments) +set -euo pipefail + +WORKSPACE="$( cd -- "$( dirname -- "${BASH_SOURCE[0]}" )/../../.." &> /dev/null && pwd )" +DB="$WORKSPACE/db/published_track.db" + +PLATFORM="" +SOURCE_FOLDER="" + +while [[ $# -gt 0 ]]; do + case "$1" in + --platform) PLATFORM="$2"; shift 2 ;; + --source-folder) SOURCE_FOLDER="$2"; shift 2 ;; + *) echo "未知参数: $1"; exit 1 ;; + esac +done + +if [[ -z "$PLATFORM" || -z "$SOURCE_FOLDER" ]]; then + echo "用法: query-metrics.sh --platform --source-folder " + echo " platform: wx_mp | wx_channel | xhs | zhihu | bilibili | douyin | kuaishou | toutiao | juejin | twitter | facebook | instagram | tiktok | youtube | pinterest | threads | wxwork_moments" + exit 1 +fi + +if [[ ! -f "$DB" ]]; then + echo "❌ published-track DB 不存在: $DB" + echo " 先运行 published-track 的 init-db.sh" + exit 1 +fi + +TABLE="pub_${PLATFORM}" + +# 检查表是否存在 +table_exists=$(sqlite3 "$DB" "SELECT count(*) FROM sqlite_master WHERE type='table' AND name='$TABLE';") +if [[ "$table_exists" -eq 0 ]]; then + echo "❌ 平台表不存在: $TABLE" + exit 1 +fi + +# 查询 +result=$(sqlite3 -header -column "$DB" "SELECT * FROM $TABLE WHERE source_folder='$SOURCE_FOLDER';") + +if [[ -z "$result" ]]; then + echo "⚠️ 未找到记录: platform=$PLATFORM, source_folder=$SOURCE_FOLDER" + exit 0 +fi + +echo "$result" diff --git a/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/score-and-record.sh b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/score-and-record.sh new file mode 100755 index 00000000..80e6a090 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/score-and-record.sh @@ -0,0 +1,12 @@ +#!/usr/bin/env bash +# score-and-record.sh — 已合并入 published-track/scripts/record.sh(薄 wrapper) +# +# record.sh 现在统一处理:提供 --cal-* 分数 → cal_enabled=1 + 算 composite; +# 不提供 → cal_enabled=0。本脚本保留为兼容入口,转调 record.sh。 +# +# 打分的强制门(blind sub-agent + 阈值)在发布技能流程里执行,见各发布技能 +# SKILL.md 的"打分评估"段与 published-track/SKILL.md 块一·流程 1A。 +set -euo pipefail + +echo "ℹ️ score-and-record.sh 已合并入 record.sh,本调用转调 record.sh(兼容保留)" >&2 +exec bash "$(dirname "$0")/../../published-track/scripts/record.sh" "$@" diff --git a/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/score-only.sh b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/score-only.sh new file mode 100755 index 00000000..cdda1a19 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/content-calibrator/scripts/score-only.sh @@ -0,0 +1,113 @@ +#!/usr/bin/env bash +# score-only.sh — 仅打分不记录到 DB,用于 Agent 自查 +# 输出打分结果到 stdout(JSON),不写入 published-track DB +# +# 用法: +# score-only.sh --platform --content-path +# +# Agent 调用此脚本时,应同时传入打分参数(由 Agent LLM 打分后传入): +# score-only.sh --platform wx_mp --content-path output_articles/xxx/article.md \ +# --cal-er 3 --cal-hp 4 --cal-sr 3 --cal-ql 3 --cal-na 2 --cal-ab 4 --cal-pv 3 +set -euo pipefail + +ROOT="$(cd "$(dirname "$0")/../../.." && pwd)" +CAL_ROOT="$ROOT/calibration" + +PLATFORM="" CONTENT_PATH="" +CAL_ER="" CAL_HP="" CAL_SR="" CAL_QL="" CAL_NA="" CAL_AB="" CAL_PV="" + +while [[ $# -gt 0 ]]; do + case "$1" in + --platform) PLATFORM="$2"; shift 2 ;; + --content-path) CONTENT_PATH="$2"; shift 2 ;; + --cal-er) CAL_ER="$2"; shift 2 ;; + --cal-hp) CAL_HP="$2"; shift 2 ;; + --cal-sr) CAL_SR="$2"; shift 2 ;; + --cal-ql) CAL_QL="$2"; shift 2 ;; + --cal-na) CAL_NA="$2"; shift 2 ;; + --cal-ab) CAL_AB="$2"; shift 2 ;; + --cal-pv) CAL_PV="$2"; shift 2 ;; + *) echo "{\"ok\":false,\"error\":\"unknown arg: $1\"}"; exit 1 ;; + esac +done + +if [ -z "$PLATFORM" ]; then + echo '{"ok":false,"error":"--platform is required"}' + exit 1 +fi + +# 检查该平台是否启用 calibrator +CAL_DIR="$CAL_ROOT/$PLATFORM" +if [ ! -d "$CAL_DIR" ] || [ ! -f "$CAL_DIR/.cheat-state.json" ]; then + echo "{\"ok\":false,\"error\":\"platform $PLATFORM has content-calibrator disabled. Enable with cal-toggle.sh --platform $PLATFORM --enable\"}" + exit 1 +fi + +RUBRIC_VERSION=$(python3 -c "import json; print(json.load(open('$CAL_DIR/.cheat-state.json'))['rubric_version'])" 2>/dev/null || echo "v0") +SCORE_THRESHOLD=$(python3 -c "import json; print(json.load(open('$CAL_DIR/.cheat-state.json')).get('score_threshold',0))" 2>/dev/null || echo 0) + +# 验证打分参数 +HAS_SCORES=0 +for dim in ER HP SR QL NA AB PV; do + var_name="CAL_$dim" + if [[ -n "${!var_name}" ]]; then + HAS_SCORES=1 + val="${!var_name}" + if [[ "$val" -lt 0 || "$val" -gt 5 ]] 2>/dev/null; then + echo "{\"ok\":false,\"error\":\"cal_score_$dim=$val out of range (must be 0-5 integer)\"}" + exit 1 + fi + fi +done + +if [ "$HAS_SCORES" -eq 0 ]; then + echo '{"ok":false,"error":"no scores provided. Agent must score the content (ER/HP/SR/QL/NA/AB/PV, each 0-5) and pass via --cal-er etc."}' + exit 1 +fi + +# 算 composite +er="${CAL_ER:-0}" hp="${CAL_HP:-0}" sr="${CAL_SR:-0}" +ql="${CAL_QL:-0}" na="${CAL_NA:-0}" ab="${CAL_AB:-0}" pv="${CAL_PV:-0}" + +COMPOSITE=$(python3 -c " +er=$er; hp=$hp; sr=$sr; ql=$ql; na=$na; ab=$ab; pv=$pv +composite = (er*1.5 + hp*1.5 + sr*1.5 + ql + na + ab + pv) / 8.5 * 2.0 +print(f'{composite:.2f}') +") + +# 找最弱维度 +declare -A WEIGHTS=([ER]=1.5 [HP]=1.5 [SR]=1.5 [QL]=1.0 [NA]=1.0 [AB]=1.0 [PV]=1.0) +declare -A SCORES=([ER]="$er" [HP]="$hp" [SR]="$sr" [QL]="$ql" [NA]="$na" [AB]="$ab" [PV]="$pv") + +worst_dim="" +worst_contrib=999 +for dim in ER HP SR QL NA AB PV; do + contrib=$(python3 -c "print(${SCORES[$dim]} * ${WEIGHTS[$dim]})") + if python3 -c "exit(0 if $contrib < $worst_contrib else 1)"; then + worst_contrib=$contrib + worst_dim=$dim + fi +done + +# 输出 JSON(不写 DB)+ 阈值门判定 +python3 -c " +import json +scores = {'ER': $er, 'HP': $hp, 'SR': $sr, 'QL': $ql, 'NA': $na, 'AB': $ab, 'PV': $pv} +threshold = $SCORE_THRESHOLD +failing = [d for d, v in scores.items() if v <= threshold] +result = { + 'ok': True, + 'action': 'score_only', + 'platform': '$PLATFORM', + 'rubric_version': '$RUBRIC_VERSION', + 'scores': scores, + 'composite': $COMPOSITE, + 'worst_dim': '$worst_dim', + 'worst_contrib': $worst_contrib, + 'score_threshold': threshold, + 'passed': len(failing) == 0, + 'failing_dims': failing, + 'recorded': False +} +print(json.dumps(result, ensure_ascii=False)) +" diff --git a/addons/officials/crew/selfmedia-operator/skills/de-mouth/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/de-mouth/SKILL.md new file mode 100644 index 00000000..153e1420 --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/de-mouth/SKILL.md @@ -0,0 +1,242 @@ +--- +name: de-mouth +description: 口播视频去口误。自动识别并删除静音、语气词、卡顿词、重复句、残句等,输出干净视频+字幕+剪映草稿。触发词:去口误、剪口播、de-mouth、去除口误 +metadata: + openclaw: + emoji: ✂️ + primaryEnv: VOLCENGINE_API_KEY + requires: + bin: + - python3 + - ffmpeg +--- + +# de-mouth — 口播视频去口误 + +> 全自动口播精修:转录 → 口误检测 → 剪辑 → 输出 + +## 快速使用 + +``` +用户: 帮我把这个视频的口误剪掉 +用户: 去口误 video.mp4 +用户: 处理一下这个口播视频 +``` + +## 输出目录 + +``` +output_videos// +├── subtitles_words.json # 字级别字幕(含静音标记) +├── readable.txt # 易读格式(供 AI 分析) +├── sentences.txt # 分句列表(供 AI 分析) +├── auto_selected.json # 删除索引列表 +├── analysis.json # 分析统计 +├── _clean.mp4 # 去口误视频 +├── _clean_hd.mp4 # 高清化视频(--hd 时) +├── .srt # SRT 字幕(--srt 时) +└── jianying_draft/ # 剪映草稿目录(--draft 时) + ├── draft_content.json + └── draft_info.json +``` + +## 流程 + +``` +0. 确认视频路径 + 输出目录 + ↓ +1. 运行去口误脚本(脚本完成步骤 1-6) + ↓ +2. AI 语义分析口误(agent 执行步骤 7) + ↓ +3. 合并 AI 结果,重新剪辑 + ↓ +4. 输出最终视频 + 字幕 + 剪映草稿 +``` + +## 执行步骤 + +### 步骤 0: 确认参数 + +从用户消息中提取视频路径。确认输出目录: + +```bash +VIDEO_PATH="<用户提供的视频路径>" +VIDEO_NAME=$(basename "$VIDEO_PATH" | sed 's/\.[^.]*$//') +OUT_DIR="output_videos/${VIDEO_NAME}" +``` + +### 步骤 1: 运行去口误脚本 + +```bash +python3 ./skills/de-mouth/scripts/de_mouth.py "$VIDEO_PATH" \ + --out-dir "$OUT_DIR" \ + --srt --draft +``` + +**参数说明**: + +| 参数 | 默认值 | 说明 | +|------|--------|------| +| `--silence-threshold` | 0.3 | 静音阈值(秒) | +| `--keep-fillers` | 空 | 保留的语气词(逗号分隔,如 `嗯,啊`) | +| `--no-ai` | 关 | 跳过 AI 语义分析(只做脚本检测) | +| `--hd` | 关 | 2-pass 高清化输出 | +| `--hd-multiplier` | 1.2 | 高清化码率倍率 | +| `--srt` | 关 | 生成 SRT 字幕 | +| `--draft` | 关 | 生成剪映草稿目录 | +| `--dict` | 无 | 热词词典文件路径 | + +脚本会自动完成: +- 音频提取 → ASR 转录 → 脚本确定性检测 → 剪辑 → 输出 + +脚本完成后,读取 `analysis.json` 确认结果。 + +### 步骤 2: AI 语义分析口误 + +> 🚨 **核心原则:删前保后。所有重复/口误,删前面的,保后面的。** + +读取 `readable.txt` 和 `sentences.txt`,按以下 4 类规则分析。 + +#### 2.1 句间重复 + +**规则**:相邻句子(被静音≥0.5s 分隔)开头≥5字相同 → 删**前句整句**。 + +隔一句也要比对(中间可能是残句)。多次重复(≥3次)保留最后完整的,前面全删。 + +**输出格式**: +``` +| 句号 | idx范围 | 内容摘要 | 处理 | +|------|---------|----------|------| +| 5 | 212-233 | 与句6重复,句6更完整 | 删前句 | +``` + +#### 2.2 句内重复 + +**规则**:同一句内短语 A 出现两次(中间夹杂 1-3 字),即 A+中间+A 模式。 + +**只删前面的重复片段,不删整句。** + +``` +"于是很把于是很容易把它理解成一种不够友好但很高效的界面" + ↑删这4字↑ ↑保留后面完整内容↑ +``` + +**不是口误的情况**:列举(任务1任务2任务3)、强调(一个一个地) + +#### 2.3 残句 + +**规则**:话说到一半突然停住,后面接了静音或重新开始。 + +**整句删除**(从句首到句尾),不只是删结尾几个字。 + +判断标准: +1. 句子不完整:缺少宾语、谓语或结尾不自然 +2. 后接静音:残句后通常有明显停顿 +3. 后有重说:重新开始说类似内容 + +#### 2.4 重说纠正 + +**规则**:说错后立即纠正,删前面错误的部分。 + +| 类型 | 原文 | 删除 | +|------|------|------| +| 部分重复 | 你再关你关掉 | "你再关" | +| 否定纠正 | 它是它不是 | "它是" | +| 词被打断 | 依赖[静]依赖关系 | "依赖[静]" | + +#### 2.5 合并 AI 结果 + +将 AI 分析返回的所有删除 idx 追加到 `auto_selected.json`,去重排序。 + +**⚠️ 关键警告:行号 ≠ idx** + +``` +readable.txt 格式: idx|内容|时间 + ↑ 用这个值 + +行号1500 → "1568|[静1.02s]|..." ← idx是1568,不是1500! +``` + +**范围整段删除规则**:标记口误时,从 startIdx 到 endIdx 之间的**所有元素**(含中间的 gap)全部加入删除列表。 + +### 步骤 3: 重新剪辑(合并 AI 结果后) + +如果 AI 分析新增了删除项,需要重新剪辑: + +```bash +# 读取合并后的 auto_selected.json,转换为 delete_segments.json +# 然后调用脚本重新剪辑 +python3 ./skills/de-mouth/scripts/de_mouth.py "$VIDEO_PATH" \ + --out-dir "$OUT_DIR" \ + --apply-ai \ + --srt --draft +``` + +> 注:`--apply-ai` 模式下,脚本读取已有的 `auto_selected.json`(含 AI 追加的 idx), +> 跳过转录和检测,直接执行剪辑。 + +### 步骤 4: 输出结果 + +向用户报告: + +``` +✅ 去口误完成! + +📹 视频: output_videos//_clean.mp4 + 原时长: 19:02 → 新时长: 15:47(删除 3:15,17.1%) + +📊 检测统计: + - 静音: 114 处 + - 语气词: 89 处 + - 卡顿词: 23 处 + - 句间重复: 15 处 + - 句内重复: 8 处 + - 残句: 6 处 + - 重说纠正: 4 处 + +📄 SRT 字幕: output_videos//.srt +🎬 剪映草稿: output_videos//jianying_draft/ + (复制到 ~/Movies/JianyingPro/User Data/Projects/com.lveditor.draft/ 并重启剪映即可导入) +``` + +## ASR 说明 + +| 模式 | 条件 | 时间戳精度 | +|------|------|-----------| +| 火山引擎 | `VOLCENGINE_API_KEY` 已设置 | 字级别(毫秒精度) | +| SiliconFlow | 仅 `SILICONFLOW_API_KEY` 已设置 | 粗估(字符均匀分布) | + +火山引擎为推荐方案,提供字级别精确时间戳 + 热词词典支持。 + +## 剪映草稿说明 + +输出的 `jianying_draft/` 目录包含 `draft_content.json` + `draft_info.json`,是剪映工程的逆向格式。 + +**导入方法**: +1. 复制整个目录到 `~/Movies/JianyingPro/User Data/Projects/com.lveditor.draft/` +2. 退出剪映(Cmd+Q) +3. 重新打开剪映 +4. 首页即可看到新草稿 + +**不依赖剪映安装** — 纯文件输出,剪映未安装也不影响去口误功能。 + +## 配置 + +### 环境变量 + +| 变量 | 必需 | 说明 | +|------|------|------| +| `VOLCENGINE_API_KEY` | 推荐 | 火山引擎 ASR API Key | +| `SILICONFLOW_API_KEY` | 降级 | SiliconFlow ASR API Key | + +### 热词词典 + +可选的 `词典.txt` 文件,每行一个词,用于 ASR 转录时纠错专业术语: + +``` +Claude Code +MCP +API +openclaw +``` diff --git a/addons/officials/crew/selfmedia-operator/skills/de-mouth/scripts/de_mouth.py b/addons/officials/crew/selfmedia-operator/skills/de-mouth/scripts/de_mouth.py new file mode 100755 index 00000000..5e7dcf3b --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/de-mouth/scripts/de_mouth.py @@ -0,0 +1,1111 @@ +#!/usr/bin/env python3 +"""de-mouth — Remove filler words, silences, and speech errors from talking-head videos. + +Pipeline: + 1. Extract audio via ffmpeg + 2. Transcribe via ASR (Volcengine with word-level timestamps, or SiliconFlow fallback) + 3. Detect speech errors (script-based: silences, fillers, stutters) + 4. Output analysis files for AI semantic analysis (repetitions, corrections, incomplete sentences) + 5. Apply delete list and cut video via ffmpeg filter_complex + 6. Optionally: 2-pass HD re-encode, SRT subtitles, JianYing draft directory + +Usage: + python3 ./skills/de-mouth/scripts/de_mouth.py --out-dir [options] +""" + +import argparse +import json +import mimetypes +import os +import re +import shutil +import subprocess +import sys +import tempfile +import time +import urllib.error +import urllib.request +import uuid +from pathlib import Path + +# ── Constants ──────────────────────────────────────────────────────────────── + +ASR_VOLCENGINE_URL_SUBMIT = "https://openspeech.bytedance.com/api/v1/vc/submit" +ASR_VOLCENGINE_URL_QUERY = "https://openspeech.bytedance.com/api/v1/vc/query" +ASR_SILICONFLOW_URL = "https://api.siliconflow.cn/v1/audio/transcriptions" +UPLOAD_URL = "https://uguu.se/upload" + +DEFAULT_SILENCE_THRESHOLD = 0.3 # seconds +DEFAULT_KEEP_FILLERS = "" # comma-separated filler words to preserve +FILLER_WORDS = frozenset("嗯啊哎诶呃额唉哦噢呀欸") +STUTTER_PATTERNS = ["那个那个", "就是就是", "然后然后", "这个这个", "所以所以"] +CONTINUOUS_FILLER_PAIRS = True # detect consecutive filler pairs (嗯啊, 啊呃) + +SAFE_OUTPUT_DIRS = (Path("output_videos"), Path("tmp")) + +# ── Utilities ──────────────────────────────────────────────────────────────── + +def die(msg: str) -> None: + print(f"[error] {msg}", file=sys.stderr) + sys.exit(1) + + +def log(tag: str, msg: str) -> None: + print(f"[{tag}] {msg}") + + +def run_cmd(cmd: list[str], timeout: int = 120, check: bool = True) -> subprocess.CompletedProcess: + """Run a command, capturing output. Returns CompletedProcess.""" + try: + return subprocess.run(cmd, capture_output=True, text=True, timeout=timeout) + except subprocess.TimeoutExpired: + if check: + die(f"Command timed out: {' '.join(cmd[:3])}...") + raise + + +def ensure_safe_output_dir(raw_path: str) -> Path: + path = Path(raw_path) + if path.is_absolute(): + die("output path must be relative to the workspace") + if ".." in path.parts: + die("output path must not contain '..'") + resolved = (Path.cwd() / path).resolve() + for base in SAFE_OUTPUT_DIRS: + base_resolved = (Path.cwd() / base).resolve() + if resolved == base_resolved or resolved.is_relative_to(base_resolved): + return resolved + allowed = ", ".join(str(d) for d in SAFE_OUTPUT_DIRS) + die(f"output path must be under one of: {allowed}") + + +# ── Step 1: Audio extraction ──────────────────────────────────────────────── + +def extract_audio(video_path: str, output_path: str) -> str: + """Extract audio as MP3 from video.""" + cmd = ["ffmpeg", "-y", "-i", video_path, "-vn", "-acodec", "libmp3lame", output_path] + result = run_cmd(cmd, timeout=180) + if result.returncode != 0: + die(f"Audio extraction failed: {result.stderr[-500:]}") + if not os.path.exists(output_path): + die(f"Audio file not created: {output_path}") + log("audio", f"Extracted {os.path.getsize(output_path) / 1024 / 1024:.1f}MB MP3") + return output_path + + +# ── Step 2a: Upload audio to get public URL ───────────────────────────────── + +def upload_audio(audio_path: str) -> str: + """Upload audio to uguu.se and return public URL.""" + cmd = ["curl", "-s", "-F", f"files[]=@{audio_path}", UPLOAD_URL] + result = run_cmd(cmd, timeout=120) + if result.returncode != 0: + die(f"Upload failed: {result.stderr}") + try: + resp = json.loads(result.stdout) + url = resp["files"][0]["url"] + log("upload", f"Audio uploaded: {url}") + return url + except (json.JSONDecodeError, KeyError, IndexError) as e: + die(f"Upload response parse failed: {result.stdout[:200]}") + + +# ── Step 2b: ASR — Volcengine (word-level timestamps) ─────────────────────── + +def transcribe_volcengine(audio_url: str, api_key: str, hot_words: list[str] | None = None) -> dict: + """Transcribe via Volcengine OpenSpeech API (async submit + poll).""" + # Build request body + body = {"url": audio_url} + if hot_words: + body["hot_words"] = hot_words + + # Submit task + submit_body = json.dumps(body).encode() + req = urllib.request.Request( + f"{ASR_VOLCENGINE_URL_SUBMIT}?language=zh-CN&use_itn=True&use_capitalize=True&max_lines=1&words_per_line=15", + data=submit_body, + method="POST", + ) + req.add_header("Content-Type", "application/json") + req.add_header("x-api-key", api_key) + + try: + with urllib.request.urlopen(req, timeout=30) as resp: + submit_result = json.loads(resp.read().decode()) + except urllib.error.HTTPError as e: + die(f"Volcengine submit failed (HTTP {e.code}): {e.read().decode(errors='replace')[:300]}") + + task_id = submit_result.get("id") + if not task_id: + die(f"Volcengine submit returned no task ID: {json.dumps(submit_result)[:200]}") + + log("asr", f"Volcengine task submitted: {task_id}") + + # Poll for result + max_attempts = 120 # 10 min at 5s intervals + for attempt in range(max_attempts): + time.sleep(5) + query_req = urllib.request.Request( + f"{ASR_VOLCENGINE_URL_QUERY}?id={task_id}", + method="GET", + ) + query_req.add_header("x-api-key", api_key) + + try: + with urllib.request.urlopen(query_req, timeout=30) as resp: + query_result = json.loads(resp.read().decode()) + except urllib.error.HTTPError as e: + die(f"Volcengine query failed (HTTP {e.code})") + + code = query_result.get("code", -1) + if code == 0: + utterances = query_result.get("utterances", []) + log("asr", f"Volcengine transcription complete: {len(utterances)} utterances") + return query_result + elif code == 1000: + if attempt % 6 == 5: # log every 30s + log("asr", f"Still processing... ({attempt * 5}s)") + else: + die(f"Volcengine transcription failed (code={code})") + + die("Volcengine transcription timed out (10 min)") + + +# ── Step 2c: ASR — SiliconFlow (text only, no timestamps) ────────────────── + +def build_multipart_formdata(file_path: str, fields: dict[str, str]) -> tuple[bytes, str]: + boundary = f"----DeMouth{uuid.uuid4().hex}" + filename = os.path.basename(file_path) + content_type = mimetypes.guess_type(filename)[0] or "application/octet-stream" + parts: list[bytes] = [] + parts.append( + ( + f"--{boundary}\r\n" + f'Content-Disposition: form-data; name="file"; filename="{filename}"\r\n' + f"Content-Type: {content_type}\r\n\r\n" + ).encode("utf-8") + ) + with open(file_path, "rb") as f: + parts.append(f.read()) + parts.append(b"\r\n") + for name, value in fields.items(): + parts.append( + ( + f"--{boundary}\r\n" + f'Content-Disposition: form-data; name="{name}"\r\n\r\n' + f"{value}\r\n" + ).encode("utf-8") + ) + parts.append(f"--{boundary}--\r\n".encode("utf-8")) + return b"".join(parts), f"multipart/form-data; boundary={boundary}" + + +def transcribe_siliconflow(audio_path: str, api_key: str) -> dict: + """Transcribe via SiliconFlow API (OpenAI-compatible, text only).""" + model = os.environ.get("ASR_MODEL", "FunAudioLLM/SenseVoiceSmall") + body, content_type = build_multipart_formdata(audio_path, {"model": model}) + req = urllib.request.Request(ASR_SILICONFLOW_URL, data=body, method="POST") + req.add_header("Authorization", f"Bearer {api_key}") + req.add_header("Content-Type", content_type) + + try: + with urllib.request.urlopen(req, timeout=120) as resp: + result = json.loads(resp.read().decode()) + except urllib.error.HTTPError as e: + err_body = e.read().decode(errors="replace") + die(f"SiliconFlow ASR failed (HTTP {e.code}): {err_body[:300]}") + + text = result.get("text", "") + log("asr", f"SiliconFlow transcription complete: {len(text)} chars") + return result + + +# ── Step 2d: Unified ASR dispatch ─────────────────────────────────────────── + +def detect_asr_mode() -> str: + """Detect which ASR to use based on available API keys.""" + if os.environ.get("VOLCENGINE_API_KEY", "").strip(): + return "volcengine" + if os.environ.get("SILICONFLOW_API_KEY", "").strip(): + return "siliconflow" + die("No ASR API key found. Set VOLCENGINE_API_KEY (recommended) or SILICONFLOW_API_KEY") + + +def run_asr(audio_path: str, hot_words: list[str] | None = None) -> tuple[str, dict]: + """Run ASR and return (mode, raw_result).""" + mode = detect_asr_mode() + if mode == "volcengine": + api_key = os.environ["VOLCENGINE_API_KEY"].strip() + audio_url = upload_audio(audio_path) + result = transcribe_volcengine(audio_url, api_key, hot_words) + return "volcengine", result + else: + api_key = os.environ["SILICONFLOW_API_KEY"].strip() + result = transcribe_siliconflow(audio_path, api_key) + return "siliconflow", result + + +# ── Step 3: Generate subtitles_words.json from ASR result ─────────────────── + +def volcengine_to_words(result: dict) -> list[dict]: + """Convert Volcengine result to subtitles_words.json format.""" + all_words = [] + for utterance in result.get("utterances", []): + for word in utterance.get("words", []): + all_words.append({ + "text": word["text"], + "start": word["start_time"] / 1000, + "end": word["end_time"] / 1000, + "isGap": False, + }) + return insert_gaps(all_words) + + +def siliconflow_to_words(result: dict, audio_path: str) -> list[dict]: + """Convert SiliconFlow text result to estimated subtitles_words.json format. + + Since SiliconFlow doesn't provide timestamps, we estimate based on + character distribution across audio duration. + """ + text = result.get("text", "") + if not text: + die("SiliconFlow returned empty text") + + # Get audio duration + probe = run_cmd(["ffprobe", "-v", "quiet", "-print_format", "json", + "-show_format", audio_path], timeout=15) + try: + duration = float(json.loads(probe.stdout)["format"]["duration"]) + except (json.JSONDecodeError, KeyError, ValueError): + die("Cannot determine audio duration for timestamp estimation") + + # Split into sentences and distribute across duration + sentences = re.split(r"([。!?;\n])", text) + chunks = [] + current = "" + for part in sentences: + current += part + if part in "。!?;\n" and current.strip(): + chunks.append(current.strip()) + current = "" + if current.strip(): + chunks.append(current.strip()) + + if not chunks: + chunks = [text] + + total_chars = sum(len(c) for c in chunks) + all_words = [] + cursor = 0.0 + + for chunk in chunks: + chunk_chars = len(chunk) + chunk_duration = (chunk_chars / total_chars) * duration if total_chars > 0 else 0 + char_duration = chunk_duration / chunk_chars if chunk_chars > 0 else 0 + + for char in chunk: + if char.strip(): # skip whitespace + all_words.append({ + "text": char, + "start": round(cursor, 3), + "end": round(cursor + char_duration, 3), + "isGap": False, + }) + cursor += char_duration + + return insert_gaps(all_words) + + +def insert_gaps(words: list[dict]) -> list[dict]: + """Insert isGap entries between words where silence exists.""" + result = [] + last_end = 0.0 + + for word in words: + gap_duration = word["start"] - last_end + if gap_duration > 0.1: + if gap_duration > 0.5: + # Split long gaps into 1s chunks + gap_start = last_end + while gap_start < word["start"]: + gap_end = min(gap_start + 1.0, word["start"]) + result.append({ + "text": "", + "start": round(gap_start, 3), + "end": round(gap_end, 3), + "isGap": True, + }) + gap_start = gap_end + else: + result.append({ + "text": "", + "start": round(last_end, 3), + "end": round(word["start"], 3), + "isGap": True, + }) + result.append(word) + last_end = word["end"] + + return result + + +# ── Step 4: Script-based deterministic detection ──────────────────────────── + +def detect_silences(words: list[dict], threshold: float) -> list[int]: + """Detect silences >= threshold. Returns indices to delete.""" + indices = [] + for i, w in enumerate(words): + if w.get("isGap") and (w["end"] - w["start"]) >= threshold: + indices.append(i) + return indices + + +def detect_filler_words(words: list[dict], keep: set[str]) -> list[int]: + """Detect standalone filler words. Returns indices to delete.""" + indices = [] + for i, w in enumerate(words): + if not w.get("isGap") and w["text"] in FILLER_WORDS and w["text"] not in keep: + indices.append(i) + return indices + + +def detect_stutters(words: list[dict]) -> list[int]: + """Detect stutter patterns (那个那个, 就是就是, etc.). Returns indices to delete. + + Strategy: delete the first occurrence, keep the last. + """ + indices = [] + # Build full text with indices + indexed_text = [] + for i, w in enumerate(words): + if not w.get("isGap"): + indexed_text.append((i, w["text"])) + + # Join text for pattern matching + full_text = "".join(t for _, t in indexed_text) + idx_map = [i for i, _ in indexed_text] + + for pattern in STUTTER_PATTERNS: + half = pattern[:len(pattern)//2] + pos = 0 + while True: + idx = full_text.find(pattern, pos) + if idx == -1: + break + # Delete the first half of the stutter + half_len = len(half) + for j in range(half_len): + if idx + j < len(idx_map): + indices.append(idx_map[idx + j]) + pos = idx + len(pattern) + + return indices + + +def detect_continuous_fillers(words: list[dict], keep: set[str]) -> list[int]: + """Detect consecutive filler word pairs (嗯啊, 啊呃). Delete both.""" + indices = [] + for i in range(len(words) - 1): + w1, w2 = words[i], words[i + 1] + if (not w1.get("isGap") and not w2.get("isGap") + and w1["text"] in FILLER_WORDS and w2["text"] in FILLER_WORDS + and w1["text"] not in keep and w2["text"] not in keep): + indices.append(i) + indices.append(i + 1) + return indices + + +# ── Step 5: Generate analysis files for AI semantic analysis ──────────────── + +def generate_readable(words: list[dict], output_path: str) -> None: + """Generate readable.txt for AI analysis.""" + lines = [] + for i, w in enumerate(words): + if w.get("isGap"): + dur = (w["end"] - w["start"]) + if dur >= 0.2: + lines.append(f"{i}|[静{dur:.2f}s]|{w['start']:.2f}-{w['end']:.2f}") + else: + lines.append(f"{i}|{w['text']}|{w['start']:.2f}-{w['end']:.2f}") + Path(output_path).write_text("\n".join(lines), encoding="utf-8") + + +def generate_sentences(words: list[dict], output_path: str) -> None: + """Generate sentences.txt — split by silences >= 0.5s.""" + sentences = [] + curr = {"text": "", "startIdx": -1, "endIdx": -1} + + for i, w in enumerate(words): + is_long_gap = w.get("isGap") and (w["end"] - w["start"]) >= 0.5 + if is_long_gap: + if curr["text"]: + sentences.append(curr) + curr = {"text": "", "startIdx": -1, "endIdx": -1} + elif not w.get("isGap"): + if curr["startIdx"] == -1: + curr["startIdx"] = i + curr["text"] += w["text"] + curr["endIdx"] = i + + if curr["text"]: + sentences.append(curr) + + lines = [] + for i, s in enumerate(sentences): + lines.append(f"{i}|{s['startIdx']}-{s['endIdx']}|{s['text']}") + Path(output_path).write_text("\n".join(lines), encoding="utf-8") + return sentences + + +# ── Step 6: FFmpeg precise cutting ────────────────────────────────────────── + +def probe_video(filepath: str) -> dict: + """Probe video parameters.""" + result = run_cmd(["ffprobe", "-v", "quiet", "-print_format", "json", + "-show_format", "-show_streams", filepath], timeout=15) + if result.returncode != 0: + die(f"ffprobe failed: {result.stderr}") + + data = json.loads(result.stdout) + video_stream = next((s for s in data.get("streams", []) if s.get("codec_type") == "video"), None) + if not video_stream: + die("No video stream found") + + format_info = data.get("format", {}) + duration = float(format_info.get("duration", 0)) + + bitrate = int(video_stream.get("bit_rate", 0)) + if bitrate == 0: + bitrate = int(format_info.get("bit_rate", 0)) + + profile = video_stream.get("profile", "high") + pix_fmt = video_stream.get("pix_fmt", "yuv420p") + + return { + "duration": duration, + "bitrate": bitrate, + "profile": profile, + "pix_fmt": pix_fmt, + } + + +def delete_indices_to_segments(words: list[dict], delete_indices: list[int]) -> list[dict]: + """Convert delete indices to time segments. + + For each deleted word, expand to include adjacent gaps. + Then merge overlapping/adjacent segments. + """ + if not delete_indices: + return [] + + # Mark all indices to delete, expanding to include adjacent gaps + delete_set = set(delete_indices) + + # Expand: if a word is deleted, also delete any gap between it and adjacent deleted words + sorted_indices = sorted(delete_set) + expanded = set(sorted_indices) + + # Merge contiguous ranges and convert to time segments + segments = [] + idx_list = sorted(expanded) + + i = 0 + while i < len(idx_list): + start_idx = idx_list[i] + end_idx = idx_list[i] + + # Extend range while contiguous or with gaps in between + j = i + 1 + while j < len(idx_list): + # Check if there are only gaps between end_idx and idx_list[j] + gap_only = True + for k in range(end_idx + 1, idx_list[j]): + if k < len(words) and not words[k].get("isGap"): + gap_only = False + break + if gap_only and idx_list[j] - end_idx <= 3: # allow up to 2 gaps between + end_idx = idx_list[j] + j += 1 + else: + break + + start_time = words[start_idx]["start"] + end_time = words[end_idx]["end"] + segments.append({"start": round(start_time, 3), "end": round(end_time, 3)}) + i = j + + # Merge overlapping segments + segments.sort(key=lambda s: s["start"]) + merged = [] + for seg in segments: + if merged and seg["start"] <= merged[-1]["end"] + 0.2: # 200ms merge gap + merged[-1]["end"] = max(merged[-1]["end"], seg["end"]) + else: + merged.append({"start": seg["start"], "end": seg["end"]}) + + return merged + + +def cut_video(input_path: str, delete_segments: list[dict], output_path: str) -> None: + """Cut video using ffmpeg -ss/-to per segment + concat demuxer. + + This avoids trim filter grey frame issues and provides frame-accurate cutting. + """ + if not delete_segments: + # No deletions, just copy + shutil.copy2(input_path, output_path) + log("cut", "No segments to delete, copied original") + return + + info = probe_video(input_path) + duration = info["duration"] + bitrate_k = info["bitrate"] // 1000 + profile = info["profile"].lower() + pix_fmt = info["pix_fmt"] + + # Map profile + x264_profile = "high" + if profile == "main": + x264_profile = "main" + elif profile == "baseline": + x264_profile = "baseline" + + maxrate_k = bitrate_k * 13 // 10 + bufsize_k = bitrate_k * 2 + + # Calculate keep segments + keep_segs = [] + cursor = 0.0 + for seg in delete_segments: + if seg["start"] > cursor: + keep_segs.append({"start": cursor, "end": seg["start"]}) + cursor = seg["end"] + if cursor < duration: + keep_segs.append({"start": cursor, "end": duration}) + + if not keep_segs: + die("All segments would be deleted") + + log("cut", f"Keeping {len(keep_segs)} segments, deleting {len(delete_segments)} segments") + + # Calculate deleted time + deleted_time = sum(s["end"] - s["start"] for s in delete_segments) + log("cut", f"Deleting {deleted_time:.2f}s of {duration:.2f}s ({deleted_time/duration*100:.1f}%)") + + # Use filter_complex for precise cutting + filters = [] + vconcat = "" + + for i, seg in enumerate(keep_segs): + filters.append( + f"[0:v]trim=start={seg['start']:.3f}:end={seg['end']:.3f},setpts=PTS-STARTPTS[v{i}]" + ) + filters.append( + f"[0:a]atrim=start={seg['start']:.3f}:end={seg['end']:.3f},asetpts=PTS-STARTPTS[a{i}]" + ) + vconcat += f"[v{i}]" + + filters.append(f"{vconcat}concat=n={len(keep_segs)}:v=1:a=0[outv]") + + # Audio: simple concat (no crossfade for speed) + aconcat = "".join(f"[a{i}]" for i in range(len(keep_segs))) + filters.append(f"{aconcat}concat=n={len(keep_segs)}:v=0:a=1[outa]") + + filter_complex = ";".join(filters) + + cmd = [ + "ffmpeg", "-y", "-v", "error", "-stats", + "-i", input_path, + "-filter_complex", filter_complex, + "-map", "[outv]", "-map", "[outa]", + "-c:v", "libx264", f"-profile:v", x264_profile, + f"-b:v", f"{bitrate_k}k", f"-maxrate", f"{maxrate_k}k", f"-bufsize", f"{bufsize_k}k", + "-pix_fmt", pix_fmt, + "-c:a", "aac", "-b:a", "192k", + "-movflags", "+faststart", + output_path, + ] + + result = run_cmd(cmd, timeout=600) + if result.returncode != 0: + # Fallback: segment-by-segment cutting + log("cut", "filter_complex failed, falling back to segment extraction...") + _cut_video_fallback(input_path, keep_segs, output_path, x264_profile, bitrate_k, maxrate_k, bufsize_k, pix_fmt) + else: + log("cut", f"Output: {output_path}") + + +def _cut_video_fallback(input_path: str, keep_segs: list[dict], output_path: str, + profile: str, bitrate_k: int, maxrate_k: int, bufsize_k: int, + pix_fmt: str) -> None: + """Fallback: extract each segment with -ss/-to, then concat.""" + tmp_dir = tempfile.mkdtemp(prefix="de_mouth_") + try: + part_files = [] + for i, seg in enumerate(keep_segs): + part_file = os.path.join(tmp_dir, f"part{i:04d}.mp4") + seg_duration = seg["end"] - seg["start"] + cmd = [ + "ffmpeg", "-y", + "-ss", f"{seg['start']:.3f}", "-i", input_path, + "-t", f"{seg_duration:.3f}", + "-c:v", "libx264", f"-profile:v", profile, + f"-b:v", f"{bitrate_k}k", f"-maxrate", f"{maxrate_k}k", f"-bufsize", f"{bufsize_k}k", + "-pix_fmt", pix_fmt, + "-c:a", "aac", "-b:a", "192k", + "-avoid_negative_ts", "make_zero", + part_file, + ] + result = run_cmd(cmd, timeout=120) + if result.returncode != 0: + die(f"Segment {i} extraction failed: {result.stderr[-300:]}") + part_files.append(part_file) + + # Concat + list_file = os.path.join(tmp_dir, "concat.txt") + list_content = "\n".join(f"file '{os.path.abspath(f)}'" for f in part_files) + Path(list_file).write_text(list_content) + + cmd = [ + "ffmpeg", "-y", "-v", "error", + "-f", "concat", "-safe", "0", "-i", list_file, + "-c", "copy", "-movflags", "+faststart", + output_path, + ] + result = run_cmd(cmd, timeout=120) + if result.returncode != 0: + die(f"Concat failed: {result.stderr[-300:]}") + + log("cut", f"Output (fallback): {output_path}") + finally: + shutil.rmtree(tmp_dir, ignore_errors=True) + + +# ── Step 7: HD re-encode (2-pass) ─────────────────────────────────────────── + +def hd_reencode(input_path: str, output_path: str, multiplier: float = 1.2) -> None: + """2-pass encode with sharpening, matching or exceeding original quality.""" + info = probe_video(input_path) + bitrate_k = info["bitrate"] // 1000 + target_k = int(bitrate_k * multiplier) + maxrate_k = target_k * 13 // 10 + bufsize_k = target_k * 2 + profile = info["profile"].lower() + pix_fmt = info["pix_fmt"] + + x264_profile = "high" + if profile == "main": + x264_profile = "main" + elif profile == "baseline": + x264_profile = "baseline" + + passlog = tempfile.mktemp(prefix="ffmpeg2pass_") + try: + # Pass 1 + cmd = [ + "ffmpeg", "-y", "-v", "error", + "-i", input_path, + "-vf", "unsharp=5:5:0.3:5:5:0.3", + "-c:v", "libx264", f"-profile:v", x264_profile, + f"-b:v", f"{target_k}k", "-preset", "slow", + "-pix_fmt", pix_fmt, + "-pass", "1", "-passlogfile", passlog, + "-an", "-f", "null", "/dev/null", + ] + result = run_cmd(cmd, timeout=600) + if result.returncode != 0: + die(f"HD Pass 1 failed: {result.stderr[-300:]}") + + # Pass 2 + cmd = [ + "ffmpeg", "-y", "-v", "error", "-stats", + "-i", input_path, + "-vf", "unsharp=5:5:0.3:5:5:0.3", + "-c:v", "libx264", f"-profile:v", x264_profile, + f"-b:v", f"{target_k}k", f"-maxrate", f"{maxrate_k}k", f"-bufsize", f"{bufsize_k}k", + "-preset", "slow", + "-pix_fmt", pix_fmt, + "-pass", "2", "-passlogfile", passlog, + "-c:a", "copy", + "-movflags", "+faststart", + output_path, + ] + result = run_cmd(cmd, timeout=600) + if result.returncode != 0: + die(f"HD Pass 2 failed: {result.stderr[-300:]}") + + log("hd", f"HD output: {output_path} ({bitrate_k}kbps → {target_k}kbps)") + finally: + for ext in ("", ".log", ".log.mbtree"): + try: + os.unlink(f"{passlog}{ext}") + except OSError: + pass + + +# ── Step 8: SRT subtitle generation ───────────────────────────────────────── + +def generate_srt(words: list[dict], delete_indices: set[int], output_path: str) -> None: + """Generate SRT subtitle file from words, excluding deleted indices.""" + def to_srt_time(sec: float) -> str: + h = int(sec // 3600) + m = int((sec % 3600) // 60) + s = int(sec % 60) + ms = int(round((sec % 1) * 1000)) + return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}" + + # Group words into subtitle lines (by sentence boundaries) + lines = [] + current_words = [] + sentence_end_re = re.compile(r"[。!?!?;]") + + for i, w in enumerate(words): + if i in delete_indices: + continue + if w.get("isGap"): + if current_words: + gap_dur = w["end"] - w["start"] + if gap_dur >= 0.3: # sentence break at 0.3s+ silence + lines.append(current_words) + current_words = [] + continue + + current_words.append(w) + if sentence_end_re.search(w["text"]): + lines.append(current_words) + current_words = [] + + if current_words: + lines.append(current_words) + + # Write SRT + srt_content = [] + for idx, line_words in enumerate(lines, 1): + text = "".join(w["text"] for w in line_words).strip() + if not text: + continue + # Remove trailing punctuation for cleaner subtitles + text = re.sub(r"[。!?!?;]+$", "", text) + start = to_srt_time(line_words[0]["start"]) + end = to_srt_time(line_words[-1]["end"]) + srt_content.append(f"{idx}\n{start} --> {end}\n{text}\n") + + Path(output_path).write_text("\n".join(srt_content), encoding="utf-8") + log("srt", f"Generated {len(srt_content)} subtitle lines") + + +# ── Step 9: JianYing draft generation ─────────────────────────────────────── + +def generate_jianying_draft(srt_path: str, output_dir: str, name: str = "字幕草稿", + width: int = 1440, height: int = 1080, + font_size: int = 7, text_color: str = "#FFDE00", + border_color: str = "#000000") -> None: + """Generate JianYing draft directory (draft_content.json + draft_info.json). + + Users can copy this directory to ~/Movies/JianyingPro/User Data/Projects/com.lveditor.draft/ + and restart JianYing to import. + + This is a simplified implementation based on capcut-mate's reverse-engineered schema. + Only covers basic subtitles without effects/animations. + """ + draft_id = uuid.uuid4().hex[:16] + draft_dir = os.path.join(output_dir, f"{name}-{draft_id[-8:]}") + os.makedirs(draft_dir, exist_ok=True) + + # Parse SRT + srt_text = Path(srt_path).read_text(encoding="utf-8") + blocks = re.split(r"\n\n+", srt_text.strip()) + captions = [] + for block in blocks: + lines = block.strip().split("\n") + if len(lines) < 3: + continue + m = re.match( + r"(\d+):(\d+):(\d+)[,.](\d+)\s*-->\s*(\d+):(\d+):(\d+)[,.](\d+)", + lines[1] + ) + if not m: + continue + + def to_us(h, mi, s, ms): + return (int(h) * 3600 + int(mi) * 60 + int(s)) * 1_000_000 + int(ms) * 1000 + + start_us = to_us(*m.groups()[:4]) + end_us = to_us(*m.groups()[4:]) + text = "\n".join(lines[2:]) + captions.append({"start": start_us, "end": end_us, "text": text}) + + if not captions: + log("draft", "No captions to generate draft") + return + + # Generate minimal draft_content.json + # This is a simplified version — enough for JianYing to load subtitles + total_duration_us = captions[-1]["end"] if captions else 10_000_000 + + # Build tracks + text_materials = {} + text_tracks = [] + + for i, cap in enumerate(captions): + mat_id = uuid.uuid4().hex + text_materials[mat_id] = { + "type": "text", + "content": cap["text"], + "font_size": font_size, + "font_color": text_color, + "border_color": border_color, + "border_width": 40, + "bold": True, + "has_shadow": True, + } + + text_tracks.append({ + "id": uuid.uuid4().hex, + "material_id": mat_id, + "target_timerange": { + "start": cap["start"], + "duration": cap["end"] - cap["start"], + }, + "source_timerange": { + "start": 0, + "duration": cap["end"] - cap["start"], + }, + "transform": {"y": -850}, + }) + + draft_content = { + "id": draft_id, + "platform": "win", + "type": "video", + "duration": total_duration_us, + "materials": { + "texts": text_materials, + }, + "tracks": [ + { + "type": "text", + "attribute": 0, + "segments": text_tracks, + } + ], + "config": { + "width": width, + "height": height, + "fps": 30.0, + }, + } + + draft_info = { + "draft_id": draft_id, + "draft_name": f"{name}-{draft_id[-8:]}", + "duration": total_duration_us, + "create_time": int(time.time() * 1000000), + "update_time": int(time.time() * 1000000), + "platform": "win", + } + + content_path = os.path.join(draft_dir, "draft_content.json") + info_path = os.path.join(draft_dir, "draft_info.json") + + Path(content_path).write_text( + json.dumps(draft_content, ensure_ascii=False, indent=2), encoding="utf-8" + ) + Path(info_path).write_text( + json.dumps(draft_info, ensure_ascii=False, indent=2), encoding="utf-8" + ) + + log("draft", f"JianYing draft: {draft_dir} ({len(captions)} captions)") + log("draft", "Copy to ~/Movies/JianyingPro/User Data/Projects/com.lveditor.draft/ and restart JianYing to import") + + +# ── Main ──────────────────────────────────────────────────────────────────── + +def main() -> None: + parser = argparse.ArgumentParser( + description="Remove filler words and speech errors from talking-head videos" + ) + parser.add_argument("video_path", help="Path to the source video file") + parser.add_argument("--out-dir", required=True, dest="out_dir", + help="Output directory under output_videos/ or tmp/") + parser.add_argument("--silence-threshold", type=float, default=DEFAULT_SILENCE_THRESHOLD, + dest="silence_threshold", + help=f"Silence threshold in seconds (default: {DEFAULT_SILENCE_THRESHOLD})") + parser.add_argument("--keep-fillers", type=str, default=DEFAULT_KEEP_FILLERS, + dest="keep_fillers", + help="Comma-separated filler words to preserve (e.g. 嗯,啊)") + parser.add_argument("--no-ai", action="store_true", + help="Skip AI semantic analysis (only script-based detection)") + parser.add_argument("--hd", action="store_true", + help="2-pass HD re-encode output") + parser.add_argument("--hd-multiplier", type=float, default=1.2, + dest="hd_multiplier", + help="HD bitrate multiplier (default: 1.2)") + parser.add_argument("--srt", action="store_true", + help="Generate SRT subtitle file") + parser.add_argument("--draft", action="store_true", + help="Generate JianYing draft directory") + parser.add_argument("--dict", type=str, default=None, + help="Path to hot words dictionary file (one word per line)") + args = parser.parse_args() + + if not os.path.isfile(args.video_path): + die(f"Video file not found: {args.video_path}") + + out_dir = ensure_safe_output_dir(args.out_dir) + out_dir.mkdir(parents=True, exist_ok=True) + + # Parse keep-fillers + keep_fillers = set() + if args.keep_fillers: + keep_fillers = {w.strip() for w in args.keep_fillers.split(",") if w.strip()} + + # Load hot words dictionary + hot_words = None + if args.dict and os.path.isfile(args.dict): + hot_words = [line.strip() for line in Path(args.dict).read_text(encoding="utf-8").splitlines() if line.strip()] + log("dict", f"Loaded {len(hot_words)} hot words") + + # ── Step 1: Extract audio ── + log("step", "1/8: Extracting audio...") + tmp_dir = str(out_dir / "_tmp") + os.makedirs(tmp_dir, exist_ok=True) + + try: + audio_path = os.path.join(tmp_dir, "audio.mp3") + extract_audio(args.video_path, audio_path) + + # ── Step 2: ASR transcription ── + log("step", "2/8: Transcribing via ASR...") + asr_mode, asr_result = run_asr(audio_path, hot_words) + + # ── Step 3: Generate subtitles_words.json ── + log("step", "3/8: Generating word-level subtitles...") + if asr_mode == "volcengine": + words = volcengine_to_words(asr_result) + else: + words = siliconflow_to_words(asr_result, audio_path) + + words_path = str(out_dir / "subtitles_words.json") + Path(words_path).write_text(json.dumps(words, ensure_ascii=False, indent=2), encoding="utf-8") + log("words", f"{len(words)} elements ({sum(1 for w in words if w.get('isGap'))} gaps)") + + # ── Step 4: Script-based deterministic detection ── + log("step", "4/8: Detecting speech errors (script-based)...") + + # Silences + silence_indices = detect_silences(words, args.silence_threshold) + log("detect", f"Silences >= {args.silence_threshold}s: {len(silence_indices)}") + + # Filler words + filler_indices = detect_filler_words(words, keep_fillers) + log("detect", f"Filler words: {len(filler_indices)}") + + # Stutters + stutter_indices = detect_stutters(words) + log("detect", f"Stutters: {len(stutter_indices)}") + + # Continuous fillers + continuous_indices = detect_continuous_fillers(words, keep_fillers) + log("detect", f"Continuous fillers: {len(continuous_indices)}") + + # Merge all script-based deletions + script_deletes = sorted(set(silence_indices + filler_indices + stutter_indices + continuous_indices)) + log("detect", f"Total script-based deletions: {len(script_deletes)}") + + # ── Step 5: Generate analysis files for AI ── + log("step", "5/8: Generating analysis files...") + + readable_path = str(out_dir / "readable.txt") + sentences_path = str(out_dir / "sentences.txt") + auto_selected_path = str(out_dir / "auto_selected.json") + + generate_readable(words, readable_path) + generate_sentences(words, sentences_path) + + # Save script-based auto_selected for AI to extend + Path(auto_selected_path).write_text( + json.dumps(script_deletes, indent=2), encoding="utf-8" + ) + + # Generate analysis report + analysis = { + "mode": asr_mode, + "total_words": len(words), + "gaps": sum(1 for w in words if w.get("isGap")), + "script_detections": { + "silences": len(silence_indices), + "fillers": len(filler_indices), + "stutters": len(stutter_indices), + "continuous_fillers": len(continuous_indices), + }, + "script_delete_count": len(script_deletes), + "ai_analysis_needed": not args.no_ai, + } + analysis_path = str(out_dir / "analysis.json") + Path(analysis_path).write_text( + json.dumps(analysis, ensure_ascii=False, indent=2), encoding="utf-8" + ) + + # ── Step 6: Apply deletions and cut video ── + log("step", "6/8: Cutting video...") + + # Use script_deletes as the final delete list + # (AI analysis results will be merged by the agent and re-run if needed) + delete_segments = delete_indices_to_segments(words, script_deletes) + + video_name = Path(args.video_path).stem + cut_output = str(out_dir / f"{video_name}_clean.mp4") + cut_video(args.video_path, delete_segments, cut_output) + + # ── Step 7: HD re-encode (optional) ── + final_output = cut_output + if args.hd: + log("step", "7/8: HD re-encoding...") + hd_output = str(out_dir / f"{video_name}_clean_hd.mp4") + hd_reencode(cut_output, hd_output, args.hd_multiplier) + final_output = hd_output + else: + log("step", "7/8: HD re-encode skipped") + + # ── Step 8: SRT + JianYing draft (optional) ── + if args.srt or args.draft: + log("step", "8/8: Generating subtitles...") + srt_output = str(out_dir / f"{video_name}.srt") + delete_set = set(script_deletes) + generate_srt(words, delete_set, srt_output) + + if args.draft: + draft_dir = str(out_dir / "jianying_draft") + os.makedirs(draft_dir, exist_ok=True) + generate_jianying_draft(srt_output, draft_dir, name=video_name) + else: + log("step", "8/8: Subtitles skipped") + + # ── Summary ── + info = probe_video(args.video_path) + new_info = probe_video(final_output) + + print("\n" + "=" * 60) + print(f"✅ de-mouth complete!") + print(f" Input: {args.video_path} ({info['duration']:.1f}s)") + print(f" Output: {final_output} ({new_info['duration']:.1f}s)") + deleted_dur = info["duration"] - new_info["duration"] + print(f" Deleted: {deleted_dur:.1f}s ({deleted_dur / info['duration'] * 100:.1f}%)") + print(f" ASR: {asr_mode}") + print(f" Script detections: {len(script_deletes)} items") + if not args.no_ai: + print(f" ⚠️ AI semantic analysis pending — agent should read readable.txt + sentences.txt") + print(f" Then update auto_selected.json and re-run with --apply-ai") + if args.srt: + print(f" SRT: {out_dir / f'{video_name}.srt'}") + if args.draft: + print(f" JianYing draft: {out_dir / 'jianying_draft'}") + print("=" * 60) + + finally: + shutil.rmtree(tmp_dir, ignore_errors=True) + + +if __name__ == "__main__": + main() diff --git a/addons/officials/crew/selfmedia-operator/skills/douyin-publish/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/douyin-publish/SKILL.md index 673b37e6..917d7e90 100644 --- a/addons/officials/crew/selfmedia-operator/skills/douyin-publish/SKILL.md +++ b/addons/officials/crew/selfmedia-operator/skills/douyin-publish/SKILL.md @@ -1,8 +1,9 @@ --- name: douyin-publish -description: Publish videos to Douyin (抖音) via open platform API with OAuth2 authentication. - Supports video upload, cover image, hashtags, and privacy settings. Requires Douyin - open platform OAuth2 credentials. +description: Publish content to Douyin (抖音) via open platform H5 Schema. + Generates a schema URL to open Douyin app's publish page. Supports video, + images, album, hashtags, privacy, note mode, and forward-to-daily. Requires + Douyin open platform web app credentials. metadata: openclaw: emoji: 🎤 @@ -11,17 +12,20 @@ metadata: - python3 --- -# 抖音视频发布(douyin-publish) +# 抖音内容发布(douyin-publish) -通过抖音开放平台 API 发布视频,支持视频上传、封面、话题标签。使用 OAuth2 认证。 +通过抖音开放平台 H5 Schema 发布内容(视频/图片/图集),生成 schema URL 唤起抖音 App 完成发布。 --- ## 前置条件 -1. 抖音开放平台创建应用,获取 client_key / client_secret -2. 申请「视频管理」权限范围 -3. 首次运行需浏览器授权,后续自动使用 refresh token +1. 在 [抖音开放平台](https://open.douyin.com) 创建**网页应用**,获取 `client_key` / `client_secret` +2. 申请 `h5.share` 能力权限(H5 场景分享/发布) +3. 投稿能力(抖音 30.5.0+):额外申请 `aweme.share` 权限 +4. 转发到日常能力(抖音 30.5.0+):额外申请 `aweme.forward` 权限 + +> **H5 Schema 方式不需要用户 OAuth2 授权**,使用 client_token(client_credential 授予)即可生成 schema。用户在手机端打开 schema 后由抖音 App 处理发布。 --- @@ -32,13 +36,10 @@ metadata: ```json { "client_key": "your_client_key", - "client_secret": "your_client_secret", - "redirect_uri": "https://localhost:8080/callback" + "client_secret": "your_client_secret" } ``` -首次授权后 token 保存到 `~/.openclaw/credentials/douyin_token.json` - --- ## 使用方式 @@ -46,34 +47,55 @@ metadata: ```bash python3 ./skills/douyin-publish/scripts/publish_douyin.py \ --title "视频标题" \ - --video video.mp4 \ - --cover cover.jpg \ + --video "https://example.com/video.mp4" \ --tags 话题1,话题2 ``` +图集模式: + +```bash +python3 ./skills/douyin-publish/scripts/publish_douyin.py \ + --title "图文标题" \ + --images "https://example.com/img1.jpg,https://example.com/img2.jpg" \ + --tags 话题1 +``` + --- ## 参数说明 | 参数 | 必填 | 说明 | |------|------|------| -| `--title` | 是 | 视频标题,最多 55 字 | -| `--video` | 是 | 视频文件路径,支持 mp4,建议 9:16 | -| `--cover` | 否 | 封面图 URL 或本地路径 | +| `--title` | 是 | 内容标题 | +| `--video` | 否* | 视频 URL(公网可访问,mp4/mov,≤128M) | +| `--image` | 否* | 单张图片 URL(png/jpg/gif,≤20M) | +| `--images` | 否* | 逗号分隔的图片 URL(图集模式,png/jpg,抖音 22.2.0+) | | `--tags` | 否 | 逗号分隔的话题标签 | -| `--private` | 否 | 设为仅自己可见 | +| `--short-title` | 否 | 短标题(抖音 30.0.0+) | +| `--private-status` | 否 | 可见范围:0=公开,1=仅自己,2=好友可见(抖音 30.0.0+) | +| `--download-type` | 否 | 下载控制:1=允许,2=不允许(抖音 30.0.0+) | +| `--share-to-type` | 否 | 发布类型:0=投稿,1=转发到日常(抖音 25.4.0+) | +| `--poi-id` | 否 | 地理位置 POI ID(抖音 22.2.0+) | +| `--feature` | 否 | 设为 `note` 启用笔记模式(抖音 30.3.0+,仅多图) | + +*\*`--video`、`--image`、`--images` 至少提供一个。`--video` 优先于图片参数。* + +> **重要**:`--video` / `--image` / `--images` 均为**公网可访问的 URL**,不是本地文件路径。需先将媒体文件上传到可公网访问的位置。 --- ## Agent 工作流 -1. 检查抖音配置和 token -2. 准备视频 + 标题 +1. 检查抖音配置(`douyin_config.json`) +2. 确保视频/图片已上传到公网可访问的 URL 3. 运行 `publish_douyin.py` 脚本 4. 检查 stdout JSON 输出: - - `{"ok": true, "item_id": "xxx", "url": "https://www.douyin.com/video/xxx"}` → 成功 - - `{"ok": false, "error": "AUTH_REQUIRED"}` → 需要完成 OAuth2 授权 + - `{"ok": true, "schema_url": "snssdk1128://...", "share_id": "xxx"}` → 成功生成 schema + - `{"ok": false, "error": "CONFIG_MISSING"}` → 需要配置凭据 - `{"ok": false, "error": "..."}` → 其他错误 +5. 将 `schema_url` 提供给用户,用户在手机上打开(扫码或点击链接) +6. 用户在抖音 App 中确认发布 +7. 可通过 `share_id` 调用「查询视频分享结果」API 获取发布状态 --- @@ -81,26 +103,9 @@ python3 ./skills/douyin-publish/scripts/publish_douyin.py \ | 错误 | 原因 | 处理 | |------|------|------| -| AUTH_REQUIRED | 无有效 OAuth2 token | 提示用户完成授权 | -| UPLOAD_FAILED | 上传失败 | 检查文件格式,重试一次 | -| PUBLISH_FAILED | 发布失败 | 检查权限和内容合规性 | -| QUOTA_EXCEEDED | API 调用频率限制 | 等待后重试 | - -## 发布记录(强制) - -发布成功后,**必须**立即调用 `published-track` 技能记录发布信息: - -```bash -./skills/published-track/scripts/record.sh \ - --platform douyin \ - --title "标题" \ - --content-type video \ - --source-folder "<原始文件夹路径>" \ - --publish-url "<发布URL>" \ - --publish-date "$(date +%Y-%m-%d)" -``` - -`--source-folder` 为原始内容所在的相对路径(如 `output_articles/xxx` 或 `output_videos/xxx`)。 -`--publish-url` 为发布后获得的 URL,若发布失败则留空并在 `--notes` 中注明原因。 - -执行 `./skills/published-track/scripts/init-db.sh`(幂等,重复执行无副作用)。 +| CONFIG_MISSING | 无 douyin_config.json | 创建配置文件 | +| CONFIG_INVALID | 缺少 client_key 或 client_secret | 补全配置 | +| CLIENT_TOKEN_FAILED | 获取 client_token 失败 | 检查凭据是否正确、应用是否审核通过 | +| TICKET_FAILED | 获取 ticket 失败 | 检查应用是否审核通过、h5.share 权限是否已申请 | +| SHARE_ID_FAILED | 获取 share_id 失败 | 检查 h5.share 权限 | +| MISSING_MEDIA | 未提供视频或图片 | 至少提供一种媒体内容 | diff --git a/addons/officials/crew/selfmedia-operator/skills/douyin-publish/scripts/publish_douyin.py b/addons/officials/crew/selfmedia-operator/skills/douyin-publish/scripts/publish_douyin.py index 4ec96953..4e566b99 100755 --- a/addons/officials/crew/selfmedia-operator/skills/douyin-publish/scripts/publish_douyin.py +++ b/addons/officials/crew/selfmedia-operator/skills/douyin-publish/scripts/publish_douyin.py @@ -1,20 +1,28 @@ #!/usr/bin/env python3 -"""Publish videos to Douyin via open platform API with OAuth2.""" +"""Publish content to Douyin via H5 Schema (open platform). + +Generates a schema URL that opens the Douyin app's publish page. +The user must open this URL on a device with the Douyin app installed +to confirm and complete the publishing. + +Flow: + client_key + client_secret → client_token → ticket + share_id → schema URL +""" import argparse import hashlib import json -import os +import random +import string import sys import time -import uuid from pathlib import Path +from urllib.parse import quote, urlencode import requests CREDS_DIR = Path.home() / ".openclaw" / "credentials" CONFIG_FILE = CREDS_DIR / "douyin_config.json" -TOKEN_FILE = CREDS_DIR / "douyin_token.json" DOUYIN_API = "https://open.douyin.com" @@ -30,176 +38,288 @@ def err_exit(msg: str, code: int = 1) -> None: def load_config() -> dict: if not CONFIG_FILE.exists(): - err_exit("AUTH_REQUIRED: no douyin_config.json", 2) - return json.loads(CONFIG_FILE.read_text()) + err_exit( + "CONFIG_MISSING: no douyin_config.json. " + "Create ~/.openclaw/credentials/douyin_config.json with client_key and client_secret.", + 2, + ) + cfg = json.loads(CONFIG_FILE.read_text()) + if not cfg.get("client_key") or not cfg.get("client_secret"): + err_exit("CONFIG_INVALID: douyin_config.json must contain client_key and client_secret", 2) + return cfg -def load_token() -> dict: - if not TOKEN_FILE.exists(): - err_exit("AUTH_REQUIRED", 2) - return json.loads(TOKEN_FILE.read_text()) +def generate_nonce_str(length: int = 32) -> str: + chars = string.ascii_letters + string.digits + return "".join(random.choices(chars, k=length)) -def refresh_access_token(config: dict, token_data: dict) -> str: - refresh_token = token_data.get("refresh_token", "") - if not refresh_token: - err_exit("AUTH_REQUIRED: no refresh token", 2) +def get_client_token(config: dict) -> str: + """Get client_token via client_credential grant (no user auth needed). + client_token is valid for 2 hours. Repeated calls invalidate the previous + one (with a 5-minute buffer). Rate limit: 500 calls per 5 minutes. + """ resp = requests.post( - f"{DOUYIN_API}/oauth/refresh_token/", - data={ + f"{DOUYIN_API}/oauth/client_token/", + json={ "client_key": config["client_key"], "client_secret": config["client_secret"], - "grant_type": "refresh_token", - "refresh_token": refresh_token, + "grant_type": "client_credential", }, - headers={"Content-Type": "application/x-www-form-urlencoded"}, + headers={"Content-Type": "application/json"}, timeout=30, ) if resp.status_code != 200: - err_exit(f"AUTH_REQUIRED: refresh failed: {resp.text}", 2) + err_exit(f"CLIENT_TOKEN_FAILED: HTTP {resp.status_code}: {resp.text[:200]}") data = resp.json() - token_info = data.get("data", {}) - if token_info.get("error_code") != 0: - err_exit(f"AUTH_REQUIRED: {token_info.get('description', data)}", 2) - - new_token = { - "access_token": token_info.get("access_token", ""), - "refresh_token": token_info.get("refresh_token", refresh_token), - "expires_in": token_info.get("expires_in", 0), - "open_id": token_info.get("open_id", token_data.get("open_id", "")), - } - CREDS_DIR.mkdir(parents=True, exist_ok=True) - TOKEN_FILE.write_text(json.dumps(new_token, indent=2)) - return new_token["access_token"] - - -def upload_video(access_token: str, open_id: str, video_path: str) -> str: - url = f"{DOUYIN_API}/api/douyin/v1/video/upload_video/" - file_size = os.path.getsize(video_path) - filename = os.path.basename(video_path) - - with open(video_path, "rb") as f: - resp = requests.post( - url, - headers={"access-token": access_token}, - params={"open_id": open_id}, - files={"video": (filename, f, "video/mp4")}, - timeout=300, - ) + if data.get("message") != "success": + desc = data.get("data", {}).get("description", str(data)) + err_exit(f"CLIENT_TOKEN_FAILED: {desc}") - if resp.status_code in (401, 403): - err_exit("AUTH_REQUIRED", 2) - data = resp.json() - if data.get("data", {}).get("error_code") != 0: - err_exit(f"UPLOAD_FAILED: {data.get('data', {}).get('description', data)}") - - video_id = data.get("data", {}).get("video", {}).get("video_id", "") - if not video_id: - err_exit(f"UPLOAD_FAILED: no video_id: {data}") - return video_id - - -def upload_cover(access_token: str, open_id: str, cover_path: str) -> str: - url = f"{DOUYIN_API}/api/douyin/v1/video/upload_cover/" - with open(cover_path, "rb") as f: - resp = requests.post( - url, - headers={"access-token": access_token}, - params={"open_id": open_id}, - files={"image": ("cover.jpg", f, "image/jpeg")}, - timeout=60, - ) - data = resp.json() - return data.get("data", {}).get("image", {}).get("image_url", "") + token = data.get("data", {}).get("access_token", "") + if not token: + err_exit("CLIENT_TOKEN_FAILED: no access_token in response") + return token -def create_video( - access_token: str, open_id: str, - video_id: str, title: str, cover_url: str, - tags: list[str], is_private: bool, -) -> dict: - url = f"{DOUYIN_API}/api/douyin/v1/video/create_video/" - text = title - if tags: - text += " " + " ".join(f"#{t}" for t in tags) +def get_open_ticket(client_token: str) -> str: + """Get open ticket for schema signature generation.""" + resp = requests.get( + f"{DOUYIN_API}/open/getticket/", + headers={ + "Content-Type": "application/json", + "access-token": client_token, + }, + timeout=30, + ) + if resp.status_code != 200: + err_exit(f"TICKET_FAILED: HTTP {resp.status_code}: {resp.text[:200]}") - payload = { - "video_id": video_id, - "text": text[:55], - } - if cover_url: - payload["cover_url"] = cover_url - if is_private: - payload["private"] = True + data = resp.json() + ticket = data.get("data", {}).get("ticket", "") + if not ticket: + err_exit("TICKET_FAILED: no ticket in response") + return ticket - resp = requests.post( - url, + +def get_share_id(client_token: str) -> str: + """Get share_id for tracking publish result via webhook/query.""" + resp = requests.get( + f"{DOUYIN_API}/share-id/", + params={"need_callback": "true"}, headers={ - "access-token": access_token, "Content-Type": "application/json", + "access-token": client_token, }, - params={"open_id": open_id}, - json=payload, timeout=30, ) + if resp.status_code != 200: + err_exit(f"SHARE_ID_FAILED: HTTP {resp.status_code}: {resp.text[:200]}") - if resp.status_code in (401, 403): - err_exit("AUTH_REQUIRED", 2) data = resp.json() - if data.get("data", {}).get("error_code") != 0: - msg = data.get("data", {}).get("description", str(data)) - if "login" in msg.lower() or "登录" in msg: - err_exit("AUTH_REQUIRED", 2) - err_exit(f"PUBLISH_FAILED: {msg}") + error_code = data.get("extra", {}).get("error_code", -1) + if error_code != 0: + desc = data.get("extra", {}).get("description", str(data)) + err_exit(f"SHARE_ID_FAILED: {desc}") + + share_id = data.get("data", {}).get("share_id", "") + if not share_id: + err_exit("SHARE_ID_FAILED: no share_id in response") + return share_id + + +def generate_signature(ticket: str, nonce_str: str, timestamp: str) -> str: + """Generate MD5 signature for H5 schema. + + Sign string: nonce_str={nonce_str}&ticket={ticket}×tamp={timestamp} + Result: MD5 hex digest of the sign string. + """ + sign_str = f"nonce_str={nonce_str}&ticket={ticket}×tamp={timestamp}" + return hashlib.md5(sign_str.encode()).hexdigest() + + +def build_schema_url( + client_key: str, + ticket: str, + share_id: str, + title: str, + video_path: str | None, + image_path: str | None, + image_list_path: list[str] | None, + hashtag_list: list[str] | None, + title_hashtag_list: list[dict] | None, + short_title: str | None, + private_status: int | None, + download_type: int | None, + share_to_type: int | None, + poi_id: str | None, + feature: str | None, +) -> str: + """Build the H5 share schema URL. + + Schema format: snssdk1128://openplatform/share?share_type=h5&client_key=xx&... + All parameters should be URL-encoded. Spaces encoded as %20 (not +). + """ + nonce_str = generate_nonce_str(32) + timestamp = str(int(time.time())) + signature = generate_signature(ticket, nonce_str, timestamp) + + params = { + "client_key": client_key, + "nonce_str": nonce_str, + "timestamp": timestamp, + "signature": signature, + "share_type": "h5", + } - item_id = data.get("data", {}).get("item_id", "") - return {"ok": True, "item_id": item_id, "url": f"https://www.douyin.com/video/{item_id}"} + if share_id: + params["state"] = share_id + + # Media content + if video_path: + params["video_path"] = video_path + params["share_to_publish"] = "1" + if image_path: + params["image_path"] = image_path + if image_list_path: + params["image_list_path"] = json.dumps(image_list_path, ensure_ascii=False) + + # Title + if title: + params["title"] = title + if short_title: + params["short_title"] = short_title + + # Hashtags + if hashtag_list: + params["hashtag_list"] = json.dumps(hashtag_list, ensure_ascii=False) + if title_hashtag_list: + params["title_hashtag_list"] = json.dumps(title_hashtag_list, ensure_ascii=False) + + # Privacy & download + if private_status is not None: + params["private_status"] = str(private_status) + if download_type is not None: + params["download_type"] = str(download_type) + if share_to_type is not None: + params["share_to_type"] = str(share_to_type) + + # Location + if poi_id: + params["poi_id"] = poi_id + + # Note mode + if feature: + params["feature"] = feature + + # Use quote_via=quote so spaces become %20 instead of + + return "snssdk1128://openplatform/share?" + urlencode(params, quote_via=quote) def main() -> None: - parser = argparse.ArgumentParser(description="Publish video to Douyin") - parser.add_argument("--title", required=True, help="Video title (max 55 chars)") - parser.add_argument("--video", required=True, help="Video file path") - parser.add_argument("--cover", help="Cover image path") - parser.add_argument("--tags", default="", help="Comma-separated tags") - parser.add_argument("--private", action="store_true", help="Set to private") + parser = argparse.ArgumentParser( + description="Publish content to Douyin via H5 Schema (open platform)" + ) + parser.add_argument("--title", required=True, help="Content title") + parser.add_argument( + "--video", + help="Video URL (publicly accessible, mp4/mov, max 128M)", + ) + parser.add_argument( + "--image", + help="Single image URL (png/jpg/gif, max 20M)", + ) + parser.add_argument( + "--images", + help="Comma-separated image URLs for album mode (png/jpg, Douyin 22.2.0+)", + ) + parser.add_argument("--tags", default="", help="Comma-separated hashtags") + parser.add_argument( + "--short-title", dest="short_title", + help="Short title (Douyin 30.0.0+)", + ) + parser.add_argument( + "--private-status", dest="private_status", type=int, choices=[0, 1, 2], + help="Visibility: 0=public, 1=self-only, 2=friends-only (Douyin 30.0.0+)", + ) + parser.add_argument( + "--download-type", dest="download_type", type=int, choices=[1, 2], + help="Download: 1=allow, 2=disallow (Douyin 30.0.0+)", + ) + parser.add_argument( + "--share-to-type", dest="share_to_type", type=int, choices=[0, 1], + help="Publish type: 0=post, 1=forward to daily (Douyin 25.4.0+)", + ) + parser.add_argument( + "--poi-id", dest="poi_id", + help="Location POI ID (Douyin 22.2.0+)", + ) + parser.add_argument( + "--feature", choices=["note"], + help="Set to 'note' for note mode with multi-image (Douyin 30.3.0+)", + ) args = parser.parse_args() - if not os.path.exists(args.video): - err_exit(f"UPLOAD_FAILED: video not found: {args.video}") + # Validate: at least one media required + if not args.video and not args.image and not args.images: + err_exit("MISSING_MEDIA: at least one of --video, --image, --images is required") config = load_config() - token_data = load_token() - try: - access_token = refresh_access_token(config, token_data) - except SystemExit: - raise - except Exception as e: - err_exit(f"AUTH_REQUIRED: {e}", 2) + # Step 1: Get client_token (no user auth needed) + sys.stderr.write("[douyin-publish] step 1/4: getting client_token...\n") + client_token = get_client_token(config) - open_id = token_data.get("open_id", "") - if not open_id: - err_exit("AUTH_REQUIRED: no open_id in token", 2) + # Step 2: Get open ticket for signature + sys.stderr.write("[douyin-publish] step 2/4: getting open ticket...\n") + ticket = get_open_ticket(client_token) - sys.stderr.write("[douyin-publish] uploading video...\n") - video_id = upload_video(access_token, open_id, args.video) - - cover_url = "" - if args.cover and os.path.exists(args.cover): - sys.stderr.write("[douyin-publish] uploading cover...\n") - cover_url = upload_cover(access_token, open_id, args.cover) + # Step 3: Get share_id for result tracking + sys.stderr.write("[douyin-publish] step 3/4: getting share_id...\n") + share_id = get_share_id(client_token) + # Build hashtag lists tags = [t.strip() for t in args.tags.split(",") if t.strip()] if args.tags else [] - - sys.stderr.write("[douyin-publish] creating video post...\n") - result = create_video( - access_token, open_id, video_id, args.title, cover_url, - tags, args.private, + hashtag_list = tags if tags else None + title_hashtag_list = None + if tags: + # Place hashtags at end of title (AiToEarn convention) + title_hashtag_list = [{"name": tag, "start": len(args.title)} for tag in tags] + + # Build image list + image_list_path = None + if args.images: + image_list_path = [url.strip() for url in args.images.split(",") if url.strip()] + + # Step 4: Generate schema URL + sys.stderr.write("[douyin-publish] step 4/4: generating share schema...\n") + schema_url = build_schema_url( + client_key=config["client_key"], + ticket=ticket, + share_id=share_id, + title=args.title, + video_path=args.video, + image_path=args.image, + image_list_path=image_list_path, + hashtag_list=hashtag_list, + title_hashtag_list=title_hashtag_list, + short_title=args.short_title, + private_status=args.private_status, + download_type=args.download_type, + share_to_type=args.share_to_type, + poi_id=args.poi_id, + feature=args.feature, ) - output(result) + + sys.stderr.write(f"[douyin-publish] schema generated (share_id={share_id})\n") + output({ + "ok": True, + "schema_url": schema_url, + "share_id": share_id, + "hint": "Open schema_url on a device with Douyin app to complete publishing. Use share_id to query result later.", + }) if __name__ == "__main__": diff --git a/addons/officials/crew/selfmedia-operator/skills/facebook-publish/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/facebook-publish/SKILL.md index 716e72fc..586567c7 100644 --- a/addons/officials/crew/selfmedia-operator/skills/facebook-publish/SKILL.md +++ b/addons/officials/crew/selfmedia-operator/skills/facebook-publish/SKILL.md @@ -102,24 +102,3 @@ python3 ./skills/facebook-publish/scripts/publish_facebook.py \ | UPLOAD_FAILED | 上传失败 | 检查文件格式,重试 | | MEDIA_PROCESSING | 视频处理中 | 轮询状态等待完成 | | RATE_LIMIT | 频率限制 | 等待后重试 | - ---- - -## 发布记录(强制) - -发布成功后,**必须**立即调用 `published-track` 技能记录发布信息: - -```bash -./skills/published-track/scripts/record.sh \ - --platform facebook \ - --title "标题" \ - --content-type post \ - --source-folder "<原始文件夹路径>" \ - --publish-url "<发布URL>" \ - --publish-date "$(date +%Y-%m-%d)" -``` - -`--source-folder` 为原始内容所在的相对路径(如 `output_articles/xxx` 或 `output_videos/xxx`)。 -`--publish-url` 为发布后获得的 URL,若发布失败则留空并在 `--notes` 中注明原因。 - -执行 `./skills/published-track/scripts/init-db.sh`(幂等,重复执行无副作用)。 diff --git a/addons/officials/crew/selfmedia-operator/skills/generate-wenyan-theme/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/generate-wenyan-theme/SKILL.md new file mode 100644 index 00000000..149866ac --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/generate-wenyan-theme/SKILL.md @@ -0,0 +1,351 @@ +--- +name: generate-wenyan-theme +description: Generate custom WeChat CSS themes from natural language descriptions, + WeChat article URLs, or recent articles from a WeChat Official Account. + Produces a valid CSS file conforming to wenyan and ready for wx-mp-publisher. +metadata: + openclaw: + emoji: 🎨 + requires: + bins: + - node +--- + +# 微信公众号自定义主题 CSS 生成器 + +根据用户的自然语言需求,生成符合微信公众号排版规范的自定义 CSS 样式表,保存为本地文件。 + +--- + +## 核心能力 + +- **自然语言转 CSS**:理解视觉需求(如"赛博朋克风"、"深色代码块"、"带装饰的引用块"),转换为精确的 CSS 代码 +- **微信文章仿样式生成**:当用户提供 `https://mp.weixin.qq.com` 文章链接时,调用 `wx-mp-hunter fetch --html` 获取正文 HTML,分析原文排版特征后生成 wenyan CSS +- **公众号近期文章归纳生成**:当用户提供公众号账号时,调用 `wx-mp-hunter search` + `account-posts` + `fetch --html` 采集近期文章 HTML,抽取共性后生成模板 +- **主题注册联动**:生成自定义 CSS 后,自动更新同 crew 内 `./skills/wx-mp-publisher/SKILL.md` 的主题列表,方便后续发布优先选用 +- **微信排版规范适配**:严格遵循 `#wenyan` 命名空间约束,确保样式能完美注入微信公众号 DOM 结构 +- **高级排版特效**:支持伪元素 (`::before`/`::after`)、渐变背景 (`linear-gradient`)、内联 SVG/Base64 图片等高级 CSS 特性 + +--- + +## 输入模式识别 + +本技能支持三种模式,按以下优先级判断: + +1. **单篇文章 URL 模式**:用户输入包含 `https://mp.weixin.qq.com` 开头的链接。 +2. **公众号账号模式**:用户明确提供微信公众号账号名、别名或要求“参考某公众号/抓取某公众号最近文章”。 +3. **自然语言模式**:没有微信文章链接或公众号账号时,按普通视觉需求生成。 + +--- + +## 文章采集脚本 + +脚本路径:`./skills/generate-wenyan-theme/scripts/collect-theme-sources.js` + +调用方式:`node ./skills/generate-wenyan-theme/scripts/collect-theme-sources.js ...` + +该脚本会调用全局 `wx-mp-hunter` wrapper: + +- URL 模式:`wx-mp-hunter fetch --html` +- 账号模式:`wx-mp-hunter search ` → `account-posts ` → 对候选文章逐篇 `fetch --html` + +### URL 模式 + +```bash +node ./skills/generate-wenyan-theme/scripts/collect-theme-sources.js --url --output wenyan-theme-sources.json +``` + +输出 JSON 中 `articles[0].content_html` 为文章正文 HTML。 + +### 公众号账号模式 + +```bash +node ./skills/generate-wenyan-theme/scripts/collect-theme-sources.js --account <公众号名> --count 10 --output wenyan-theme-sources.json +``` + +如果用户同时给出关键词或筛选信息,传入 `--keywords`: + +```bash +node ./skills/generate-wenyan-theme/scripts/collect-theme-sources.js --account <公众号名> --keywords "关键词1,关键词2" --count 10 --scan-batch 20 --max-scan 100 --output wenyan-theme-sources.json +``` + +筛选规则: + +- 无关键词:默认取最近 10 篇。 +- 有关键词:先抓最近 20 篇,按标题、摘要、作者匹配关键词;不足目标数量时继续抓下一批 20 篇,直到满足数量、无更多文章或达到 `--max-scan`。 +- 每篇文章会通过 `fetch --html` 获取 `content_html`。 + +> 若 `wx-mp-hunter` 返回 `SESSION_EXPIRED`,按 `wx-mp-hunter` 技能的扫码登录流程处理后重试原命令。 + +--- + +## HTML 样式分析要点 + +从 `content_html` 中抽取共性时,只分析可迁移到 wenyan CSS 的视觉规律,不复制微信原文中不可控或依赖原始 DOM 的实现细节: + +- 颜色:正文色、标题色、强调色、引用/代码/分割线背景色 +- 字号与层级:h1/h2/h3、正文、注释、小字之间的相对比例 +- 间距节奏:段落间距、标题上下留白、列表缩进、引用块 padding +- 装饰语言:标题前后缀、底纹、边框、圆角、分割线、卡片感 +- 内容类型:是否频繁使用图片、引用、列表、代码、表格 +- 共同约束:多篇账号样本中重复出现的风格才作为模板核心;单篇偶发元素只作为可选细节 + +--- + + +### 1. 强制命名空间约束(最重要) + +所有 CSS 选择器 **必须** 以 `#wenyan` 开头,中间用空格隔开。缺少 `#wenyan` 前缀的样式将失效。 + +- ✅ `#wenyan h1 { color: red; }` +- ❌ `h1 { color: red; }` + +### 2. 字体与字号 + +- **font-family**:严禁主动设置,保持默认以适配微信公众号编辑器的系统字体 +- **font-size**:建议 12px - 18px 范围,避免排版溢出或阅读困难 + +### 3. 支持的 CSS 选择器字典 + +| 目标元素 | CSS 选择器 | 常用定制属性 | +|---------|-----------|------------| +| 全局默认 | `#wenyan` | `background-image`, `line-height`, `color` | +| 各级标题 | `#wenyan h1` ~ `#wenyan h6` | `font-size`, `text-align`, `border-bottom`, `margin` | +| 标题文字 | `#wenyan h1 span` | `color`, `font-weight`, `background` | +| 标题装饰 | `#wenyan h1::before` | `content`, `display`, `width`, `height`, `background-image` | +| 段落文本 | `#wenyan p` | `text-indent`, `letter-spacing`, `color` | +| 引用块 | `#wenyan blockquote` | `border-left`, `background-color`, `padding` | +| 代码块外层 | `#wenyan pre` | `background-color`, `border-radius`, `padding`, `overflow-x: auto` | +| 代码块内容 | `#wenyan pre code` | `color` | +| 分割线 | `#wenyan hr` | `border`, `border-top-style`, `border-color` | +| 超链接 | `#wenyan a` | `color`, `text-decoration`, `border-bottom` | + +### 4. 外部资源引用限制 + +- **禁止本地路径**:严禁 `url("./bg.png")` 等本地路径 +- **合法引入方式**: + - Data URI(推荐):`url("data:image/svg+xml;utf8,...")` + - HTTPS 地址:`url(https://example.com/bg.jpg)` +- **禁止 Web 字体**:不支持 `@font-face`,只能使用系统字体 + +### 5. 输出文件路径约束 + +- 采集输出 JSON:必须是当前工作目录下的单个 `.json` 文件名,禁止目录、绝对路径和 `..` 上跳。 +- 生成的 CSS:只写入当前工作目录内的相对 `.css` 路径,禁止绝对路径、`..` 上跳、隐藏目录和非 `.css` 后缀。 + +--- + +## 参考模板(default.css 结构) + +```css +/* 全局属性 */ +#wenyan { + line-height: 1.75; + font-size: 16px; +} + +/* 标题与段落间距 */ +#wenyan h1, +#wenyan h2, +#wenyan h3, +#wenyan h4, +#wenyan h5, +#wenyan h6, +#wenyan p { + margin: 1em 0; +} + +/* 一级标题 */ +#wenyan h1 { + text-align: center; + text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.1); + font-size: 1.5em; +} + +/* 二级标题 */ +#wenyan h2 { + text-align: center; + font-size: 1.2em; + border-bottom: 1px solid #f7f7f7; + font-weight: bold; +} + +/* 列表 */ +#wenyan > ul, +#wenyan > ol { + padding-left: 1rem; +} + +#wenyan ul, +#wenyan ol { + margin-left: 1rem; + font-size: 0.9rem; +} + +/* 图片 */ +#wenyan img { + max-width: 100%; + height: auto; + margin: 0 auto; + display: block; +} + +/* 表格 */ +#wenyan table { + border-collapse: collapse; + margin: 1.4em auto; + max-width: 100%; + table-layout: fixed; + text-align: left; + overflow: auto; + display: table; +} + +/* 引用块 */ +#wenyan blockquote { + background: #afb8c133; + border-left: 0.5em solid #ccc; + margin: 1.5em 0; + padding: 0.5em 10px; + font-style: italic; + font-size: 0.9em; +} + +/* 行内代码 */ +#wenyan p code { + color: #ff502c; + padding: 4px 6px; + font-size: 0.78em; +} + +/* 代码块外围 */ +#wenyan pre { + border-radius: 5px; + line-height: 2; + margin: 1em 0.5em; + padding: .5em; + box-shadow: rgba(0, 0, 0, 0.55) 0px 1px 5px; + font-size: 12px; +} + +/* 代码块 */ +#wenyan pre code { + display: block; + overflow-x: auto; + margin: .5em; + padding: 0; +} + +/* 分割线 */ +#wenyan hr { + border: none; + border-top: 1px solid #ddd; + margin-top: 2em; + margin-bottom: 2em; +} + +/* 链接 */ +#wenyan a { + word-wrap: break-word; + color: #0069c2; +} +``` + +--- + +## Agent 执行步骤 + +### A. 自然语言模式 + +1. **分析需求**:提取关键词(如:深色、科技风、可爱),确定主色调和风格方向 +2. **生成 CSS**:严格按照命名空间约束和上述规范,生成完整的 CSS 代码 +3. **保存文件**:将 CSS 写入本地文件(如 `custom-theme.css`) +4. **注册主题**:更新 `./skills/wx-mp-publisher/SKILL.md` 的“主题选择”表格,追加或更新该自定义主题记录 +5. **后续引导**:提示使用 `wx-mp-publisher` 技能的第二个位置参数传入自定义 CSS 路径进行发布 + +### B. 单篇文章 URL 模式 + +1. **识别链接**:确认用户输入包含 `https://mp.weixin.qq.com` 开头的文章 URL。 +2. **采集 HTML**:运行采集脚本: + ```bash + node ./skills/generate-wenyan-theme/scripts/collect-theme-sources.js --url --output wenyan-theme-sources.json + ``` +3. **分析样式**:读取输出 JSON,基于 `articles[0].content_html` 分析标题、段落、引用、分割线、强调、图片周边等样式特征。 +4. **生成 CSS**:将可迁移特征映射到 `#wenyan` 选择器体系,不复制无效的微信原始 class 或 inline style。 +5. **保存文件、注册主题并引导发布**。 + +### C. 公众号账号模式 + +1. **识别账号与筛选意图**:提取公众号账号名;如果用户提供关键词、主题、人群或文章类型,将其整理为 `--keywords`。 +2. **采集样本**: + - 无筛选信息:抓最近 10 篇。 + - 有筛选信息:从最近 20 篇开始筛选,不足则继续下一批 20 篇。 + ```bash + node ./skills/generate-wenyan-theme/scripts/collect-theme-sources.js --account <公众号名> --count 10 --output wenyan-theme-sources.json + ``` + 或: + ```bash + node ./skills/generate-wenyan-theme/scripts/collect-theme-sources.js --account <公众号名> --keywords "关键词1,关键词2" --count 10 --scan-batch 20 --max-scan 100 --output wenyan-theme-sources.json + ``` +3. **向用户确认**:生成 CSS 前,必须向用户展示拟参考的文章列表(标题、发布时间/链接、匹配关键词),并询问是否继续。用户确认后再生成。 +4. **抽取共性**:优先使用多篇文章共同出现的视觉规律;冲突样式按出现频次和标题层级一致性取舍。 +5. **生成 CSS**:输出一个适合该账号整体调性的 wenyan 主题,而不是拼贴单篇文章的局部样式。 +6. **保存文件、注册主题并引导发布**。 + + +--- + +## 生成前确认要求 + +仅公众号账号模式需要生成前确认。确认消息应包含: + +- 公众号名称/别名 +- 实际采集文章数量 +- 文章标题列表 +- 如有关键词:说明命中的关键词或筛选依据 +- 输出主题文件名建议 + +用户确认“继续/可以/确认”后,才能写 CSS 文件。 + +--- + +## 生成主题注册规则 + +`generate-wenyan-theme` 与 `wx-mp-publisher` 都是 Media Operator 的私有技能,目录相对位置固定。因此每次成功生成自定义 CSS 后,必须同步更新: + +```text +./skills/wx-mp-publisher/SKILL.md +``` + +在 `wx-mp-publisher/SKILL.md` 的“主题选择”表格中追加或更新一行自定义主题记录: + +```markdown +| `` | 用户自定义:<风格摘要>(文件:``) | 用户明确指定参考该主题时优先采用;相似内容可优先建议 | +``` + +注册要求: + +- `theme-id` 使用 CSS 文件名去掉 `.css` 后缀,例如 `custom-theme.css` → `custom-theme`。 +- CSS 文件路径写相对路径,优先使用当前 crew workspace 内路径。 +- 如果同名 `theme-id` 已存在,更新原行,不重复追加。 +- 自定义主题必须在风格描述中标注“用户自定义”。 +- 自定义主题的适用场景必须强调:**用户指定参考时优先采用**。 +- 不要修改内置主题 ID 的含义。 + +注册后,`wx-mp-publisher` 发布时仍通过第二个位置参数使用 CSS 文件: + +```bash +./skills/wx-mp-publisher/scripts/publish-wx-mp.sh article.md custom-theme.css +``` + +--- + +## 与 wx-mp-publisher 配合使用 + +生成 CSS 文件后,在发布时通过自定义主题参数引用: + +```bash +./skills/wx-mp-publisher/scripts/publish-wx-mp.sh article.md custom-theme.css +``` + +> 注:当 theme 参数指向本地 `.css` 文件路径时,wenyan-cli 会将其作为自定义主题加载。 diff --git a/addons/officials/crew/selfmedia-operator/skills/generate-wenyan-theme/scripts/collect-theme-sources.js b/addons/officials/crew/selfmedia-operator/skills/generate-wenyan-theme/scripts/collect-theme-sources.js new file mode 100644 index 00000000..52a6e0eb --- /dev/null +++ b/addons/officials/crew/selfmedia-operator/skills/generate-wenyan-theme/scripts/collect-theme-sources.js @@ -0,0 +1,319 @@ +#!/usr/bin/env node +/** + * collect-theme-sources.js — collect WeChat article HTML samples for theme generation. + */ + +import { spawn } from "node:child_process"; +import { constants } from "node:fs"; +import { access, open, stat } from "node:fs/promises"; +import { basename, dirname, join, resolve, sep } from "node:path"; +import { fileURLToPath } from "node:url"; + +const DEFAULT_OUTPUT = "wenyan-theme-sources.json"; +const DEFAULT_ACCOUNT_COUNT = 10; +const DEFAULT_SCAN_BATCH = 20; +const DEFAULT_MAX_SCAN = 100; + +function printJson(data) { + process.stdout.write(`${JSON.stringify(data, null, 2)}\n`); +} + +function fail(message, code = 1) { + printJson({ ok: false, error: message }); + process.exit(code); +} + +function usage() { + process.stdout.write( + [ + "Usage:", + " node ./skills/generate-wenyan-theme/scripts/collect-theme-sources.js --url [--output file]", + " node ./skills/generate-wenyan-theme/scripts/collect-theme-sources.js --account <公众号名> [--keywords k1,k2] [--count 10] [--output file]", + "", + "Options:", + " --scan-batch N 每批扫描文章数,默认 20", + " --max-scan N 关键词筛选最多扫描文章数,默认 100", + "", + "Output:", + " --output 必须是当前工作目录下的单个 .json 文件名,不允许目录、绝对路径或 .. 上跳。", + ].join("\n") + "\n" + ); +} + +function readFlag(args, flag) { + const idx = args.indexOf(flag); + if (idx < 0 || idx + 1 >= args.length) return null; + const value = args[idx + 1]; + return value && !value.startsWith("--") ? value : null; +} + +function readNumberFlag(args, flag, fallback) { + const raw = readFlag(args, flag); + if (!raw) return fallback; + const value = Number(raw); + return Number.isFinite(value) && value > 0 ? Math.floor(value) : fallback; +} + +function parseKeywords(raw) { + if (!raw) return []; + return raw + .split(/[,,\n]/g) + .map((item) => item.trim()) + .filter(Boolean); +} + +function defaultWxHunterPath() { + const currentFile = fileURLToPath(import.meta.url); + const officialPlusRoot = resolve(dirname(currentFile), "../../../../.."); + return join(officialPlusRoot, "skills", "wx-mp-hunter", "scripts", "wx-mp-hunter.sh"); +} + +function isWechatArticleUrl(value) { + try { + const url = new URL(value); + return url.protocol === "https:" && url.hostname === "mp.weixin.qq.com"; + } catch { + return false; + } +} + +function parseArgs() { + const args = process.argv.slice(2); + if (args.includes("--help") || args.includes("-h")) { + usage(); + process.exit(0); + } + + const url = readFlag(args, "--url") ?? ""; + const account = readFlag(args, "--account") ?? ""; + if (url && account) fail("--url 与 --account 只能二选一"); + if (!url && !account) fail("必须提供 --url 或 --account"); + if (url && !isWechatArticleUrl(url)) fail("--url 必须是 https://mp.weixin.qq.com 域名下的文章链接"); + + return { + mode: url ? "url" : "account", + url, + account, + keywords: parseKeywords(readFlag(args, "--keywords")), + count: readNumberFlag(args, "--count", DEFAULT_ACCOUNT_COUNT), + scanBatch: Math.min(readNumberFlag(args, "--scan-batch", DEFAULT_SCAN_BATCH), DEFAULT_SCAN_BATCH), + maxScan: readNumberFlag(args, "--max-scan", DEFAULT_MAX_SCAN), + output: readFlag(args, "--output") ?? DEFAULT_OUTPUT, + wxHunter: defaultWxHunterPath(), + }; +} + +async function assertExecutableFile(filePath, label) { + const info = await stat(filePath).catch(() => null); + if (!info) fail(`找不到 ${label}: ${filePath}`); + if (!info.isFile()) fail(`${label} 不是文件: ${filePath}`); + await access(filePath, constants.X_OK).catch(() => fail(`${label} 不可执行: ${filePath}`)); +} + +function workspaceOutputPath(filePath) { + if (!filePath.endsWith(".json")) fail("--output 必须使用 .json 后缀"); + if (basename(filePath) !== filePath || filePath.startsWith(sep)) { + fail("--output 必须是当前工作目录下的单个 .json 文件名,不能包含目录、绝对路径或 .. 上跳"); + } + const cwd = resolve(process.cwd()); + const absolute = resolve(cwd, filePath); + const relativePrefix = `${cwd}${sep}`; + if (absolute === cwd || !absolute.startsWith(relativePrefix)) { + fail("--output 必须位于当前工作目录内"); + } + return absolute; +} + +function runJson(command, args) { + return new Promise((resolvePromise, rejectPromise) => { + const child = spawn(command, args, { stdio: ["ignore", "pipe", "pipe"], shell: false }); + let stdout = ""; + let stderr = ""; + + child.stdout.setEncoding("utf-8"); + child.stderr.setEncoding("utf-8"); + child.stdout.on("data", (chunk) => { + stdout += chunk; + }); + child.stderr.on("data", (chunk) => { + stderr += chunk; + }); + child.on("error", rejectPromise); + child.on("close", (code) => { + let parsed = null; + try { + parsed = JSON.parse(stdout); + } catch { + rejectPromise(new Error(`命令输出不是 JSON: ${basename(command)} ${args.join(" ")}\n${stderr || stdout}`)); + return; + } + + if (code !== 0) { + const error = String(parsed.error ?? stderr ?? `命令失败,退出码 ${code}`); + rejectPromise(new Error(error)); + return; + } + resolvePromise(parsed); + }); + }); +} + +function toArticleSummary(item) { + const link = String(item.link ?? "").trim(); + if (!isWechatArticleUrl(link)) return null; + + const createTime = Number(item.create_time ?? NaN); + return { + title: String(item.title ?? "").trim(), + link, + digest: String(item.digest ?? "").trim(), + author: String(item.author ?? "").trim(), + create_time: Number.isFinite(createTime) ? createTime : null, + item_show_type: Number(item.item_show_type ?? 0), + is_deleted: Boolean(item.is_deleted), + }; +} + +function articleMatches(article, keywords) { + if (keywords.length === 0) return true; + const haystack = `${article.title}\n${article.digest}\n${article.author}`.toLowerCase(); + return keywords.some((keyword) => haystack.includes(keyword.toLowerCase())); +} + +async function searchAccount(wxHunter, keyword) { + const data = await runJson(wxHunter, ["search", keyword, "--begin", "0", "--size", "5"]); + const accounts = Array.isArray(data.accounts) ? data.accounts : []; + if (accounts.length === 0) fail(`未搜索到公众号: ${keyword}`); + return accounts[0]; +} + +async function listCandidateArticles(options, fakeid) { + const selected = []; + const seen = new Set(); + let begin = 0; + const targetCount = options.keywords.length > 0 ? options.count : Math.min(options.count, DEFAULT_ACCOUNT_COUNT); + + while (selected.length < targetCount && begin < options.maxScan) { + const data = await runJson(options.wxHunter, [ + "account-posts", + fakeid, + "--begin", + String(begin), + "--size", + String(options.scanBatch), + ]); + const rawArticles = Array.isArray(data.articles) ? data.articles : []; + if (rawArticles.length === 0) break; + + for (const raw of rawArticles) { + const article = toArticleSummary(raw); + if (!article || article.is_deleted || seen.has(article.link)) continue; + seen.add(article.link); + if (articleMatches(article, options.keywords)) selected.push(article); + if (selected.length >= targetCount) break; + } + + begin += options.scanBatch; + } + + return selected; +} + +async function fetchArticleHtml(wxHunter, article) { + const data = await runJson(wxHunter, ["fetch", article.link, "--html"]); + return { + ...article, + title: String(data.title ?? article.title), + author: String(data.author ?? article.author), + publish_time: String(data.publish_time ?? ""), + content_text: String(data.content_text ?? ""), + content_html: String(data.content_html ?? ""), + }; +} + +async function writeOutput(filePath, data) { + const absolute = workspaceOutputPath(filePath); + const file = await open(absolute, constants.O_WRONLY | constants.O_CREAT | constants.O_TRUNC | constants.O_NOFOLLOW, 0o600).catch( + (error) => { + throw new Error(`无法安全写入输出文件: ${error instanceof Error ? error.message : String(error)}`); + } + ); + try { + await file.writeFile(`${JSON.stringify(data, null, 2)}\n`, "utf-8"); + } finally { + await file.close(); + } + return absolute; +} + +async function collectByUrl(options) { + const sample = await fetchArticleHtml(options.wxHunter, { + title: "", + link: options.url, + digest: "", + author: "", + create_time: null, + item_show_type: 0, + is_deleted: false, + }); + const output = await writeOutput(options.output, { + mode: "url", + source_url: options.url, + collected_at: new Date().toISOString(), + articles: [sample], + }); + printJson({ ok: true, mode: "url", output, count: 1, articles: [{ title: sample.title, link: sample.link }] }); +} + +async function collectByAccount(options) { + const account = await searchAccount(options.wxHunter, options.account); + const fakeid = String(account.fakeid ?? ""); + if (!fakeid) fail(`公众号搜索结果缺少 fakeid: ${options.account}`); + + const candidates = await listCandidateArticles(options, fakeid); + if (candidates.length === 0) fail("未找到符合条件的文章"); + + const samples = []; + for (const article of candidates) { + samples.push(await fetchArticleHtml(options.wxHunter, article)); + } + + const output = await writeOutput(options.output, { + mode: "account", + account: { + fakeid, + nickname: String(account.nickname ?? options.account), + alias: String(account.alias ?? ""), + signature: String(account.signature ?? ""), + }, + keywords: options.keywords, + collected_at: new Date().toISOString(), + scanned_limit: options.maxScan, + articles: samples, + }); + + printJson({ + ok: true, + mode: "account", + output, + count: samples.length, + account: { nickname: String(account.nickname ?? options.account), alias: String(account.alias ?? "") }, + articles: samples.map((item) => ({ title: item.title, link: item.link, publish_time: item.publish_time })), + }); +} + +async function main() { + const options = parseArgs(); + await assertExecutableFile(options.wxHunter, "wx-mp-hunter wrapper"); + + if (options.mode === "url") { + await collectByUrl(options); + return; + } + await collectByAccount(options); +} + +main().catch((error) => { + const message = error instanceof Error ? error.message : String(error); + fail(message); +}); diff --git a/addons/officials/crew/selfmedia-operator/skills/highlight-clipper/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/highlight-clipper/SKILL.md index 20528391..fe3597b4 100644 --- a/addons/officials/crew/selfmedia-operator/skills/highlight-clipper/SKILL.md +++ b/addons/officials/crew/selfmedia-operator/skills/highlight-clipper/SKILL.md @@ -48,7 +48,7 @@ python3 ./skills/highlight-clipper/scripts/clip.py --out-dir output | `--out-dir` | — | 输出目录,必须在 `output_videos/` 或 `tmp/` 下(必需) | | `--count` | 3 | 提取高光片段数量 | | `--min-duration` | 15 | 最短片段时长(秒) | -| `--max-duration` | 60 | 最长片段时长(秒) | +| `--max-duration` | 58 | 最长片段时长(秒),默认 58 留余量防超 60s 平台限制 | | `--buffer` | 3 | 片段前后缓冲(秒) | 示例 — 从一段 5 分钟视频提取 5 个高光: @@ -116,8 +116,10 @@ cat output_videos//highlights.json - 疑问句和感叹号 — 权重 1.5 - 数据/数字出现 — 权重 1.0 - 信息密度(单位时长文字量)— 权重最高 3.0 -4. **多样性选择**:贪婪选取得分最高的 N 个片段,保证片段间至少间隔 30 秒,避免高光扎堆 -5. **视频剪辑**:ffmpeg 精确裁剪,含前后缓冲秒数 +4. **多样性选择**:贪婪选取得分最高的 N 个片段,保证片段间至少间隔 3 秒(仅去重) +5. **相邻合并**:间隔 ≤ 10 秒的高光片段自动合并为一段长片段,**不限时长**——挨着的高光连成完整段落 +6. **锚点扩展**:短片段(< max-duration)以片段为中心向前后扩展填满 max-duration;已合并的长片段不截断,保留完整内容 +7. **视频剪辑**:ffmpeg 精确裁剪 --- @@ -133,3 +135,5 @@ cat output_videos//highlights.json - 转录质量取决于语音清晰度,建议使用语音清晰的视频 - 高光评分基于文本语义分析,非视觉分析——画面精彩但无语音的片段可能被遗漏 - 片段时长受 `--min-duration` 和 `--max-duration` 控制,可根据目标平台要求调整(如抖音 15–60 秒、小红书 15–45 秒) +- 相邻高光片段(间隔 ≤ 10 秒)会自动合并,合并后不限时长——适合会议中连续精彩讨论的场景 +- 短片段会以片段为中心向前后扩展至 max-duration(默认 58 秒),确保每段内容充实 diff --git a/addons/officials/crew/selfmedia-operator/skills/highlight-clipper/scripts/clip.py b/addons/officials/crew/selfmedia-operator/skills/highlight-clipper/scripts/clip.py index 87b86579..1acccd6d 100644 --- a/addons/officials/crew/selfmedia-operator/skills/highlight-clipper/scripts/clip.py +++ b/addons/officials/crew/selfmedia-operator/skills/highlight-clipper/scripts/clip.py @@ -33,8 +33,9 @@ DEFAULT_COUNT = 3 DEFAULT_BUFFER = 3.0 DEFAULT_CLIP_MIN = 15 -DEFAULT_CLIP_MAX = 60 -MIN_HIGHLIGHT_GAP = 30 # min seconds between highlight starts +DEFAULT_CLIP_MAX = 58 +MIN_HIGHLIGHT_GAP = 3 # min seconds between highlight starts (just dedup, merging handles proximity) +MERGE_GAP = 10.0 # merge highlights within this gap (seconds) SAFE_OUTPUT_DIRS = (Path("output_videos"), Path("tmp")) @@ -299,19 +300,67 @@ def select_highlights(segments: list[dict], count: int) -> list[dict]: return selected +def merge_nearby_highlights(highlights: list[dict], gap_threshold: float = MERGE_GAP) -> list[dict]: + """Merge highlights within gap_threshold of each other. No duration limit on merged clips.""" + if not highlights: + return [] + sorted_hl = sorted(highlights, key=lambda s: s.get("start", 0)) + merged: list[dict] = [dict(sorted_hl[0])] + for seg in sorted_hl[1:]: + prev = merged[-1] + prev_end = prev.get("end", 0) + curr_start = seg.get("start", 0) + if curr_start - prev_end <= gap_threshold: + merged[-1] = { + "start": prev.get("start", 0), + "end": max(prev_end, seg.get("end", 0)), + "text": prev.get("text", "") + " " + seg.get("text", ""), + "highlight_score": max(prev.get("highlight_score", 0), seg.get("highlight_score", 0)), + } + else: + merged.append(dict(seg)) + return merged + + # ── Video clipping ────────────────────────────────────────────────────── def determine_clip_bounds(seg: dict, video_duration: float, buffer: float, clip_min: float, clip_max: float) -> tuple[float, float]: + """Determine clip bounds with expansion for short segments and no truncation for merged long ones. + + - Short segment (duration < clip_max): expand to fill clip_max, centered on segment. + - Long/merged segment (duration >= clip_max): use full range + buffer, never truncate. + """ seg_start = seg.get("start", 0) seg_end = seg.get("end", 0) seg_duration = seg_end - seg_start - target = max(seg_duration + buffer * 2, clip_min) - target = min(target, clip_max) - clip_start = max(seg_start - buffer, 0) - clip_end = min(clip_start + target, video_duration) - if clip_end - clip_start < clip_min: - clip_start = max(clip_end - target, 0) + + if seg_duration >= clip_max: + # Merged/long segment — use full range + buffer, never truncate + clip_start = max(seg_start - buffer, 0) + clip_end = min(seg_end + buffer, video_duration) + else: + # Short segment — expand to fill clip_max, centered on segment + seg_center = (seg_start + seg_end) / 2 + ideal_start = seg_center - clip_max / 2 + ideal_end = seg_center + clip_max / 2 + + if ideal_start < 0: + clip_start = 0.0 + clip_end = min(clip_max, video_duration) + elif ideal_end > video_duration: + clip_end = video_duration + clip_start = max(video_duration - clip_max, 0.0) + else: + clip_start = ideal_start + clip_end = ideal_end + + # Guarantee the segment is fully inside the clip + if clip_start > seg_start: + clip_start = seg_start + if clip_end < seg_end: + clip_end = seg_end + return round(clip_start, 2), round(clip_end, 2) @@ -389,7 +438,15 @@ def main() -> None: if not highlights: die("No suitable highlights found") - print(f"[info] Selected {len(highlights)} highlights") + print(f"[info] Selected {len(highlights)} candidate highlights") + + # Merge nearby highlights into longer clips (no duration limit) + pre_merge_count = len(highlights) + highlights = merge_nearby_highlights(highlights) + if len(highlights) < pre_merge_count: + print(f"[info] Merged into {len(highlights)} clips ({pre_merge_count - len(highlights)} merges)") + else: + print("[info] No merges needed — highlights are well separated") highlights_info = [] for i, hl in enumerate(highlights, 1): diff --git a/addons/officials/crew/selfmedia-operator/skills/instagram-publish/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/instagram-publish/SKILL.md index 2709e43f..c650ba81 100644 --- a/addons/officials/crew/selfmedia-operator/skills/instagram-publish/SKILL.md +++ b/addons/officials/crew/selfmedia-operator/skills/instagram-publish/SKILL.md @@ -91,24 +91,3 @@ python3 ./skills/instagram-publish/scripts/publish_instagram.py \ | RATE_LIMIT | API 限制 | 等待后重试 | 注意:Instagram API 发布只支持 **URL** 形式的媒体,不支持本地文件上传。需要先将图片/视频上传到可公开访问的 URL。 - ---- - -## 发布记录(强制) - -发布成功后,**必须**立即调用 `published-track` 技能记录发布信息: - -```bash -./skills/published-track/scripts/record.sh \ - --platform instagram \ - --title "标题" \ - --content-type post \ - --source-folder "<原始文件夹路径>" \ - --publish-url "<发布URL>" \ - --publish-date "$(date +%Y-%m-%d)" -``` - -`--source-folder` 为原始内容所在的相对路径(如 `output_articles/xxx` 或 `output_videos/xxx`)。 -`--publish-url` 为发布后获得的 URL,若发布失败则留空并在 `--notes` 中注明原因。 - -执行 `./skills/published-track/scripts/init-db.sh`(幂等,重复执行无副作用)。 diff --git a/addons/officials/crew/selfmedia-operator/skills/juejin-publish/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/juejin-publish/SKILL.md index 13aeb5eb..cbc24c19 100644 --- a/addons/officials/crew/selfmedia-operator/skills/juejin-publish/SKILL.md +++ b/addons/officials/crew/selfmedia-operator/skills/juejin-publish/SKILL.md @@ -123,6 +123,14 @@ browser evaluate fn="(() => { const ta = document.querySelector('.CodeMirror tex b. browser upload /tmp/openclaw/uploads/cover.jpg c. 忽略可能的超时提示 + **Patchright 1.60+ 可选方案**:用 `locator.drop()` 拖拽封面图到上传区域,更稳定: + ```javascript + const buf = fs.readFileSync('/tmp/openclaw/uploads/cover.jpg'); + await page.locator('.').drop({ + files: { name: 'cover.jpg', mimeType: 'image/jpeg', buffer: buf } + }); + ``` + 4. 填写摘要(必填 *,仅兜底路径需要手动填): evaluate 找到摘要 textarea 并 fill 摘要内容取文章前 100 字左右的核心描述 @@ -172,22 +180,3 @@ See `references/publish-options.md` for category list, tag suggestions, and cove ❌ 兜底路径忘记填摘要(*必填): → 发布按钮无响应,因为摘要为空 ``` - -## 发布记录(强制) - -发布成功后,**必须**立即调用 `published-track` 技能记录发布信息: - -```bash -./skills/published-track/scripts/record.sh \ - --platform juejin \ - --title "标题" \ - --content-type article \ - --source-folder "<原始文件夹路径>" \ - --publish-url "<发布URL>" \ - --publish-date "$(date +%Y-%m-%d)" -``` - -`--source-folder` 为原始内容所在的相对路径(如 `output_articles/xxx` 或 `output_videos/xxx`)。 -`--publish-url` 为发布后获得的 URL,若发布失败则留空并在 `--notes` 中注明原因。 - -执行 `./skills/published-track/scripts/init-db.sh`(幂等,重复执行无副作用)。 diff --git a/addons/officials/crew/selfmedia-operator/skills/pinterest-publish/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/pinterest-publish/SKILL.md index f4790fd6..c53e19b6 100644 --- a/addons/officials/crew/selfmedia-operator/skills/pinterest-publish/SKILL.md +++ b/addons/officials/crew/selfmedia-operator/skills/pinterest-publish/SKILL.md @@ -83,24 +83,3 @@ python3 ./skills/pinterest-publish/scripts/publish_pinterest.py \ | INVALID_BOARD | 看板 ID 无效 | 检查看板是否存在 | 注意:Pinterest API 仅支持 URL 形式的媒体,不支持本地文件上传。 - ---- - -## 发布记录(强制) - -发布成功后,**必须**立即调用 `published-track` 技能记录发布信息: - -```bash -./skills/published-track/scripts/record.sh \ - --platform pinterest \ - --title "标题" \ - --content-type post \ - --source-folder "<原始文件夹路径>" \ - --publish-url "<发布URL>" \ - --publish-date "$(date +%Y-%m-%d)" -``` - -`--source-folder` 为原始内容所在的相对路径(如 `output_articles/xxx` 或 `output_videos/xxx`)。 -`--publish-url` 为发布后获得的 URL,若发布失败则留空并在 `--notes` 中注明原因。 - -执行 `./skills/published-track/scripts/init-db.sh`(幂等,重复执行无副作用)。 diff --git a/addons/officials/crew/selfmedia-operator/skills/published-track/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/published-track/SKILL.md deleted file mode 100644 index 8e0bcc85..00000000 --- a/addons/officials/crew/selfmedia-operator/skills/published-track/SKILL.md +++ /dev/null @@ -1,142 +0,0 @@ ---- -name: published-track -description: 发布记录追踪。使用 SQLite 数据库记录所有平台发布内容及其互动数据,按平台分表管理。发布后必须调用本技能记录,心跳巡检时更新数据。 -metadata: - openclaw: - emoji: "📊" - requires: - bins: - - bash - - sqlite3 ---- - -# published-track — 发布记录追踪 - -统一管理所有平台(微信公众号、知乎、B站、抖音、快手、小红书、今日头条、掘金、Twitter/X、Facebook、Instagram、TikTok、YouTube、Pinterest、Threads、企业微信朋友圈)的发布记录与互动数据。 - ---- - -## 数据库位置 - -`./db/published_track.db`(相对于工作区根目录) - -初始化(幂等,可重复执行): - -```bash -./skills/published-track/scripts/init-db.sh -``` - ---- - -## 平台与表对应关系 - -| 平台 | 表名 | 内容类型 | 特有指标 | -|------|------|---------|---------| -| 微信公众号 | `pub_wx_mp` | article/video/post | reads, shares, favorites, likes, comments | -| 知乎 | `pub_zhihu` | article/post | views, upvotes, comments, favorites | -| B站 | `pub_bilibili` | video | plays, danmaku, likes, coins, favorites, shares, comments | -| 抖音 | `pub_douyin` | video | plays, likes, comments, shares, favorites | -| 快手 | `pub_kuaishou` | video | plays, likes, comments, shares | -| 小红书 | `pub_xhs` | article/video/post | views, likes, favorites, comments, shares | -| 今日头条 | `pub_toutiao` | article | impressions, reads, comments, likes | -| 掘金 | `pub_juejin` | article | views, likes, comments, favorites | -| Twitter/X | `pub_twitter` | post/video | views, likes, retweets, replies, bookmarks | -| Facebook | `pub_facebook` | post/video | reach, likes, comments, shares | -| Instagram | `pub_instagram` | post/video | reach, likes, comments, shares, saves | -| TikTok | `pub_tiktok` | video | plays, likes, comments, shares, favorites | -| YouTube | `pub_youtube` | video | views, likes, comments, shares | -| Pinterest | `pub_pinterest` | post | impressions, saves, comments | -| Threads | `pub_threads` | post | views, likes, reposts, replies | -| 企业微信朋友圈 | `pub_wxwork_moments` | post | likes, comments | - ---- - -## 表结构 - -每张表共享以下通用字段: - -| 字段 | 类型 | 说明 | -|------|------|------| -| id | INTEGER PK | 自增主键 | -| title | TEXT NOT NULL | 标题 | -| content_type | TEXT NOT NULL | article / video / post | -| source_folder | TEXT NOT NULL | 原始文件夹(如 output_articles/xxx 或 output_videos/xxx) | -| publish_url | TEXT | 发布后 URL | -| publish_date | TEXT NOT NULL | 发布日期(YYYY-MM-DD) | -| notes | TEXT | 备注 | -| created_at | TEXT | 创建时间 | -| updated_at | TEXT | 更新时间 | - -去重键:`source_folder`(同一篇内容在同一平台只记录一次) - -各平台特有的互动指标字段默认值为 0,另有 `top_comment`(主要留言摘要,TEXT)字段。 - ---- - -## 使用方式 - -### 1. 发布后立即记录(强制) - -每完成一个平台的发布,**必须**立即调用 `record.sh` 记录: - -```bash -./skills/published-track/scripts/record.sh \ - --platform \ - --title "标题" \ - --content-type \ - --source-folder "output_articles/xxx" \ - --publish-url "https://..." \ - --publish-date "$(date +%Y-%m-%d)" \ - [--notes "备注"] -``` - -`--platform` 值对应上表「表名」去掉 `pub_` 前缀,如 `wx_mp`、`zhihu`、`bilibili` 等。 - -若发布失败(如平台拒绝、超时),仍需记录,`publish_url` 留空,`notes` 中注明失败原因。 - -### 2. 数据更新(心跳巡检) - -使用 `update-metrics.sh` 更新互动数据: - -```bash -./skills/published-track/scripts/update-metrics.sh \ - --platform \ - --source-folder "output_articles/xxx" \ - --reads 1234 \ - --likes 56 \ - ... -``` - -只需传入要更新的指标字段,未传入的字段保持不变。 - -### 3. 查询 - -```bash -# 查询某平台全部记录 -./skills/published-track/scripts/query.sh --platform zhihu - -# 查询某平台最近 N 条 -./skills/published-track/scripts/query.sh --platform zhihu --limit 10 - -# 查询特定内容是否已发布到某平台 -./skills/published-track/scripts/check-published.sh \ - --platform zhihu \ - --source-folder "output_articles/xxx" - -# 查询所有平台中未发布的原始文件夹 -./skills/published-track/scripts/query.sh --unpublished -``` - -### 4. 清理低数据记录 - -发布超过 7 天、互动指标低于 300 的记录,使用 `query.sh` 查出后可删除: - -```bash -./skills/published-track/scripts/query.sh --platform zhihu --stale-days 7 --below 300 -``` - ---- - -## 与发布技能的配合 - -所有发布技能(wx-mp-publisher、sync-from-mp、bilibili-publish 等)在发布成功后必须调用 `record.sh` 记录。各技能 SKILL.md 中已标注此要求,主 agent 无需额外提醒。 diff --git a/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/check-published.sh b/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/check-published.sh deleted file mode 100755 index b181fdca..00000000 --- a/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/check-published.sh +++ /dev/null @@ -1,41 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -ROOT="$(cd "$(dirname "$0")/../../.." && pwd)" -DB="$ROOT/db/published_track.db" - -if [ ! -f "$DB" ]; then - echo '{"exists":false}' - exit 0 -fi - -PLATFORM="" SOURCE_FOLDER="" - -while [[ $# -gt 0 ]]; do - case "$1" in - --platform) PLATFORM="$2"; shift 2 ;; - --source-folder) SOURCE_FOLDER="$2"; shift 2 ;; - *) echo "{\"ok\":false,\"error\":\"unknown arg: $1\"}"; exit 1 ;; - esac -done - -if [ -z "$PLATFORM" ] || [ -z "$SOURCE_FOLDER" ]; then - echo '{"ok":false,"error":"missing required args: --platform, --source-folder"}' - exit 1 -fi - -TABLE="pub_${PLATFORM}" -VALID=$(sqlite3 "$DB" "SELECT name FROM sqlite_master WHERE type='table' AND name='$TABLE';") -if [ -z "$VALID" ]; then - echo "{\"ok\":false,\"error\":\"unknown platform: $PLATFORM\"}" - exit 1 -fi - -ROW=$(sqlite3 "$DB" "SELECT id,publish_url FROM $TABLE WHERE source_folder='${SOURCE_FOLDER//\'/\'\'}';") -if [ -z "$ROW" ]; then - echo '{"exists":false}' -else - ID=$(echo "$ROW" | cut -d'|' -f1) - URL=$(echo "$ROW" | cut -d'|' -f2) - echo "{\"exists\":true,\"id\":$ID,\"publish_url\":\"$URL\"}" -fi diff --git a/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/init-db.sh b/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/init-db.sh deleted file mode 100755 index 096badda..00000000 --- a/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/init-db.sh +++ /dev/null @@ -1,301 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -ROOT="$(cd "$(dirname "$0")/../../.." && pwd)" -DB="$ROOT/db/published_track.db" - -mkdir -p "$ROOT/db" - -sqlite3 "$DB" <<'SQL' - --- 微信公众号 -CREATE TABLE IF NOT EXISTS pub_wx_mp ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - reads INTEGER DEFAULT 0, - shares INTEGER DEFAULT 0, - favorites INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - top_comment TEXT, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- 知乎 -CREATE TABLE IF NOT EXISTS pub_zhihu ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - views INTEGER DEFAULT 0, - upvotes INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - favorites INTEGER DEFAULT 0, - top_comment TEXT, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- B站 -CREATE TABLE IF NOT EXISTS pub_bilibili ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - plays INTEGER DEFAULT 0, - danmaku INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - coins INTEGER DEFAULT 0, - favorites INTEGER DEFAULT 0, - shares INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - top_comment TEXT, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- 抖音 -CREATE TABLE IF NOT EXISTS pub_douyin ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - plays INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - shares INTEGER DEFAULT 0, - favorites INTEGER DEFAULT 0, - top_comment TEXT, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- 快手 -CREATE TABLE IF NOT EXISTS pub_kuaishou ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - plays INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - shares INTEGER DEFAULT 0, - top_comment TEXT, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- 小红书 -CREATE TABLE IF NOT EXISTS pub_xhs ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - views INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - favorites INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - shares INTEGER DEFAULT 0, - top_comment TEXT, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- 今日头条 -CREATE TABLE IF NOT EXISTS pub_toutiao ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - impressions INTEGER DEFAULT 0, - reads INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - top_comment TEXT, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- 掘金 -CREATE TABLE IF NOT EXISTS pub_juejin ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - views INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - favorites INTEGER DEFAULT 0, - top_comment TEXT, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- Twitter/X -CREATE TABLE IF NOT EXISTS pub_twitter ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - views INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - retweets INTEGER DEFAULT 0, - replies INTEGER DEFAULT 0, - bookmarks INTEGER DEFAULT 0, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- Facebook -CREATE TABLE IF NOT EXISTS pub_facebook ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - reach INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - shares INTEGER DEFAULT 0, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- Instagram -CREATE TABLE IF NOT EXISTS pub_instagram ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - reach INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - shares INTEGER DEFAULT 0, - saves INTEGER DEFAULT 0, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- TikTok -CREATE TABLE IF NOT EXISTS pub_tiktok ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - plays INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - shares INTEGER DEFAULT 0, - favorites INTEGER DEFAULT 0, - top_comment TEXT, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- YouTube -CREATE TABLE IF NOT EXISTS pub_youtube ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - views INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - shares INTEGER DEFAULT 0, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- Pinterest -CREATE TABLE IF NOT EXISTS pub_pinterest ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - impressions INTEGER DEFAULT 0, - saves INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- Threads -CREATE TABLE IF NOT EXISTS pub_threads ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - views INTEGER DEFAULT 0, - likes INTEGER DEFAULT 0, - reposts INTEGER DEFAULT 0, - replies INTEGER DEFAULT 0, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - --- 企业微信朋友圈 -CREATE TABLE IF NOT EXISTS pub_wxwork_moments ( - id INTEGER PRIMARY KEY AUTOINCREMENT, - title TEXT NOT NULL, - content_type TEXT NOT NULL CHECK(content_type IN ('article','video','post')), - source_folder TEXT NOT NULL UNIQUE, - publish_url TEXT, - publish_date TEXT NOT NULL, - likes INTEGER DEFAULT 0, - comments INTEGER DEFAULT 0, - top_comment TEXT, - notes TEXT, - created_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')), - updated_at TEXT DEFAULT (strftime('%Y-%m-%d %H:%M:%S','now','localtime')) -); - -SQL - -echo '{"ok":true,"message":"published_track.db initialized"}' diff --git a/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/query.sh b/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/query.sh deleted file mode 100755 index 9358e236..00000000 --- a/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/query.sh +++ /dev/null @@ -1,91 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -ROOT="$(cd "$(dirname "$0")/../../.." && pwd)" -DB="$ROOT/db/published_track.db" - -if [ ! -f "$DB" ]; then - echo '[]' - exit 0 -fi - -PLATFORM="" LIMIT="" UNPUBLISHED=false STALE_DAYS="" BELOW="" - -while [[ $# -gt 0 ]]; do - case "$1" in - --platform) PLATFORM="$2"; shift 2 ;; - --limit) LIMIT="$2"; shift 2 ;; - --unpublished) UNPUBLISHED=true; shift ;; - --stale-days) STALE_DAYS="$2"; shift 2 ;; - --below) BELOW="$2"; shift 2 ;; - *) echo "{\"ok\":false,\"error\":\"unknown arg: $1\"}"; exit 1 ;; - esac -done - -if [ "$UNPUBLISHED" = true ]; then - # Find source_folders in output_articles/ and output_videos/ that have no record in any platform table - TABLES=$(sqlite3 "$DB" "SELECT name FROM sqlite_master WHERE type='table' AND name LIKE 'pub_%';") - FOLDERS=$(find "$ROOT/output_articles" "$ROOT/output_videos" -mindepth 1 -maxdepth 1 -type d 2>/dev/null | sed "s|$ROOT/||" | sort) - - UNPUB_LIST="[" - FIRST=true - for F in $FOLDERS; do - FOUND=false - for T in $TABLES; do - CNT=$(sqlite3 "$DB" "SELECT COUNT(*) FROM $T WHERE source_folder='${F//\'/\'\'}';") - if [ "$CNT" -gt 0 ]; then - FOUND=true - break - fi - done - if [ "$FOUND" = false ]; then - [ "$FIRST" = true ] && FIRST=false || UNPUB_LIST+="," - UNPUB_LIST+="\"$F\"" - fi - done - UNPUB_LIST+="]" - echo "$UNPUB_LIST" - exit 0 -fi - -if [ -z "$PLATFORM" ]; then - echo '{"ok":false,"error":"--platform is required (unless --unpublished)"}' - exit 1 -fi - -TABLE="pub_${PLATFORM}" -VALID=$(sqlite3 "$DB" "SELECT name FROM sqlite_master WHERE type='table' AND name='$TABLE';") -if [ -z "$VALID" ]; then - echo "{\"ok\":false,\"error\":\"unknown platform: $PLATFORM\"}" - exit 1 -fi - -# Build query -WHERE="" -if [ -n "$STALE_DAYS" ]; then - WHERE="WHERE publish_date <= date('now','-$STALE_DAYS days')" -fi - -LIMIT_CLAUSE="" -if [ -n "$LIMIT" ]; then - LIMIT_CLAUSE="LIMIT $LIMIT" -fi - -# Query all records -ROWS=$(sqlite3 -json "$DB" "SELECT * FROM $TABLE $WHERE ORDER BY publish_date DESC $LIMIT_CLAUSE;" 2>/dev/null) - -if [ -n "$BELOW" ] && [ -n "$STALE_DAYS" ]; then - # Filter for records where all main metric columns are below threshold - # Get integer columns - INT_COLS=$(sqlite3 "$DB" "PRAGMA table_info($TABLE);" | awk -F'|' '$2 != "id" && $2 != "title" && $2 != "content_type" && $2 != "source_folder" && $2 != "publish_url" && $2 != "publish_date" && $2 != "notes" && $2 != "top_comment" && $2 != "created_at" && $2 != "updated_at" {print $2}') - - CONDS="" - for C in $INT_COLS; do - [ -n "$CONDS" ] && CONDS+=" AND " - CONDS+="$C < $BELOW" - done - - ROWS=$(sqlite3 -json "$DB" "SELECT * FROM $TABLE WHERE publish_date <= date('now','-$STALE_DAYS days') AND ($CONDS) ORDER BY publish_date DESC $LIMIT_CLAUSE;" 2>/dev/null) -fi - -echo "${ROWS:-[]}" diff --git a/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/record.sh b/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/record.sh deleted file mode 100755 index ae667049..00000000 --- a/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/record.sh +++ /dev/null @@ -1,66 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -ROOT="$(cd "$(dirname "$0")/../../.." && pwd)" -DB="$ROOT/db/published_track.db" - -# Ensure db exists -if [ ! -f "$DB" ]; then - bash "$(dirname "$0")/init-db.sh" -fi - -# Parse args -PLATFORM="" TITLE="" CONTENT_TYPE="" SOURCE_FOLDER="" PUBLISH_URL="" PUBLISH_DATE="" NOTES="" - -while [[ $# -gt 0 ]]; do - case "$1" in - --platform) PLATFORM="$2"; shift 2 ;; - --title) TITLE="$2"; shift 2 ;; - --content-type) CONTENT_TYPE="$2"; shift 2 ;; - --source-folder) SOURCE_FOLDER="$2"; shift 2 ;; - --publish-url) PUBLISH_URL="$2"; shift 2 ;; - --publish-date) PUBLISH_DATE="$2"; shift 2 ;; - --notes) NOTES="$2"; shift 2 ;; - *) echo "{\"ok\":false,\"error\":\"unknown arg: $1\"}"; exit 1 ;; - esac -done - -# Validate required args -if [ -z "$PLATFORM" ] || [ -z "$TITLE" ] || [ -z "$CONTENT_TYPE" ] || [ -z "$SOURCE_FOLDER" ] || [ -z "$PUBLISH_DATE" ]; then - echo '{"ok":false,"error":"missing required args: --platform, --title, --content-type, --source-folder, --publish-date"}' - exit 1 -fi - -# Validate platform -TABLE="pub_${PLATFORM}" -VALID=$(sqlite3 "$DB" "SELECT name FROM sqlite_master WHERE type='table' AND name='$TABLE';") -if [ -z "$VALID" ]; then - echo "{\"ok\":false,\"error\":\"unknown platform: $PLATFORM (table $TABLE not found)\"}" - exit 1 -fi - -# Validate content_type -case "$CONTENT_TYPE" in - article|video|post) ;; - *) echo "{\"ok\":false,\"error\":\"invalid content_type: $CONTENT_TYPE (must be article/video/post)\"}"; exit 1 ;; -esac - -# Check duplicate -EXISTS=$(sqlite3 "$DB" "SELECT COUNT(*) FROM $TABLE WHERE source_folder='${SOURCE_FOLDER//\'/\'\'}';") -if [ "$EXISTS" -gt 0 ]; then - # Update existing record - ESC_URL="${PUBLISH_URL//\'/\'\'}" - ESC_NOTES="${NOTES//\'/\'\'}" - sqlite3 "$DB" "UPDATE $TABLE SET publish_url='$ESC_URL', notes='$ESC_NOTES', updated_at=strftime('%Y-%m-%d %H:%M:%S','now','localtime') WHERE source_folder='${SOURCE_FOLDER//\'/\'\'}';" - ID=$(sqlite3 "$DB" "SELECT id FROM $TABLE WHERE source_folder='${SOURCE_FOLDER//\'/\'\'}';") - echo "{\"ok\":true,\"action\":\"updated\",\"id\":$ID,\"table\":\"$TABLE\"}" -else - # Insert new record - ESC_TITLE="${TITLE//\'/\'\'}" - ESC_FOLDER="${SOURCE_FOLDER//\'/\'\'}" - ESC_URL="${PUBLISH_URL//\'/\'\'}" - ESC_NOTES="${NOTES//\'/\'\'}" - sqlite3 "$DB" "INSERT INTO $TABLE (title,content_type,source_folder,publish_url,publish_date,notes) VALUES ('$ESC_TITLE','$CONTENT_TYPE','$ESC_FOLDER','$ESC_URL','$PUBLISH_DATE','$ESC_NOTES');" - ID=$(sqlite3 "$DB" "SELECT last_insert_rowid();") - echo "{\"ok\":true,\"action\":\"inserted\",\"id\":$ID,\"table\":\"$TABLE\"}" -fi diff --git a/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/update-metrics.sh b/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/update-metrics.sh deleted file mode 100755 index e586f070..00000000 --- a/addons/officials/crew/selfmedia-operator/skills/published-track/scripts/update-metrics.sh +++ /dev/null @@ -1,84 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -ROOT="$(cd "$(dirname "$0")/../../.." && pwd)" -DB="$ROOT/db/published_track.db" - -if [ ! -f "$DB" ]; then - echo '{"ok":false,"error":"database not initialized, run init-db.sh first"}' - exit 1 -fi - -# Parse args -PLATFORM="" SOURCE_FOLDER="" -declare -A METRICS - -while [[ $# -gt 0 ]]; do - case "$1" in - --platform) PLATFORM="$2"; shift 2 ;; - --source-folder) SOURCE_FOLDER="$2"; shift 2 ;; - --*=*) - KEY="${1#--}" - VAL="${1#*=}" - METRICS["$KEY"]="$VAL" - shift - ;; - --*) - KEY="${1#--}" - VAL="$2" - METRICS["$KEY"]="$VAL" - shift 2 - ;; - *) echo "{\"ok\":false,\"error\":\"unknown arg: $1\"}"; exit 1 ;; - esac -done - -if [ -z "$PLATFORM" ] || [ -z "$SOURCE_FOLDER" ]; then - echo '{"ok":false,"error":"missing required args: --platform, --source-folder"}' - exit 1 -fi - -TABLE="pub_${PLATFORM}" -VALID=$(sqlite3 "$DB" "SELECT name FROM sqlite_master WHERE type='table' AND name='$TABLE';") -if [ -z "$VALID" ]; then - echo "{\"ok\":false,\"error\":\"unknown platform: $PLATFORM\"}" - exit 1 -fi - -# Check record exists -EXISTS=$(sqlite3 "$DB" "SELECT COUNT(*) FROM $TABLE WHERE source_folder='${SOURCE_FOLDER//\'/\'\'}';") -if [ "$EXISTS" -eq 0 ]; then - echo "{\"ok\":false,\"error\":\"no record found in $TABLE for source_folder=$SOURCE_FOLDER\"}" - exit 1 -fi - -# Get valid columns for this table (exclude id, created_at) -COLS=$(sqlite3 "$DB" "PRAGMA table_info($TABLE);" | awk -F'|' '{print $2}' | grep -v -E '^(id|created_at|source_folder|content_type|title|publish_date)$' | tr '\n' ' ') - -# Build SET clause -SET_PARTS=() -for KEY in "${!METRICS[@]}"; do - # Validate column exists - if ! echo " $COLS " | grep -q " $KEY "; then - echo "{\"ok\":false,\"error\":\"column '$KEY' not found in $TABLE. Valid metric columns: $COLS\"}" - exit 1 - fi - VAL="${METRICS[$KEY]}" - # Only allow integer or text values - ESC_VAL="${VAL//\'/\'\'}" - SET_PARTS+=("$KEY='$ESC_VAL'") -done - -if [ ${#SET_PARTS[@]} -eq 0 ]; then - echo '{"ok":false,"error":"no metrics provided to update"}' - exit 1 -fi - -# Always update updated_at -SET_PARTS+=("updated_at=strftime('%Y-%m-%d %H:%M:%S','now','localtime')") - -SET_CLAUSE=$(IFS=','; echo "${SET_PARTS[*]}") - -sqlite3 "$DB" "UPDATE $TABLE SET $SET_CLAUSE WHERE source_folder='${SOURCE_FOLDER//\'/\'\'}';" - -echo "{\"ok\":true,\"table\":\"$TABLE\",\"source_folder\":\"$SOURCE_FOLDER\",\"updated_columns\":${#METRICS[@]}}" diff --git a/addons/officials/crew/selfmedia-operator/skills/t2video/SKILL.md b/addons/officials/crew/selfmedia-operator/skills/t2video/SKILL.md deleted file mode 100644 index c356d6c7..00000000 --- a/addons/officials/crew/selfmedia-operator/skills/t2video/SKILL.md +++ /dev/null @@ -1,215 +0,0 @@ ---- -name: t2video -description: 一站式短视频制作工具。整合 SiliconFlow TTS 语音合成、素材搜集(Pexels/Pixabay/AI 生成)和 FFmpeg 组装,从脚本到成品视频一步完成。无 subagent 模式,适合短小视频。 -metadata: - openclaw: - emoji: "🎥" - requires: - bins: - - python3 - - ffmpeg - - ffprobe - env: - - SILICONFLOW_API_KEY - primaryEnv: SILICONFLOW_API_KEY ---- - -# t2video — 一站式短视频制作 - -Use this skill when: -- 需要从脚本生成完整短视频(TTS + 素材 + 组装) -- 用户指定主题和已有素材,生成短视频 - -**本技能是 video-producer 的简化版**:适合短小视频制作,且只生成9:16的竖屏视频。 - -**不烧录字幕**:大部分平台支持自动生成字幕,无需手动烧录。 - ---- - -## ⚙️ 执行方式(强制) - -本技能涉及多步骤生产流程,你应该 self-spawn 一个 subagent 来执行,原因:subagent 独立上下文,不会因对话历史积累而降低输出质量。 - -你只负责跟进subagent的执行,避免它们长时间卡在某个步骤,必要时可以提供提示或调整执行策略。另外在关键节点要求它向你汇报,你检查后再让它继续执行下一步。 - ---- - -## 工作流程 - -### Step 1 — 工作区目录准备 - -在 `output_videos/` 下创建项目文件夹,如 `output_videos//`,作为 project-dir。 - -工作区结构: - -``` -/ -├── script.md # 三段式脚本 -├── tts_requirement.md # 配音需求(由 Step 3 创建) -├── artifacts/ # 产出素材(TTS音频+视频素材) -│ ├── speech.mp3 # TTS 配音 -│ ├── speech.json # TTS 元数据(含 duration) -│ └──