Skip to content

Repository files navigation

CI License: MIT Python 3.12+

Sovereign-OS

A governed AI workforce, ready to deploy on day one.
Submit a goal — Sovereign-OS plans it, approves the budget, executes with built-in workers, and delivers a cryptographically verified result.

Quick StartArchitectureFeaturesDeploymentConfigurationCustom WorkersDocs

Sovereign-OS dashboard demo


Overview

Sovereign-OS is not another chatbot wrapper or agent framework. It is an operating system for autonomous AI work: a governance layer that enforces budget, quality, and permissions before any token is spent or any task is executed.

The core contract is simple. One YAML file — the Charter — declares the agent's mission, spending limits, KPIs, and allowed capabilities. Everything else — planning, approval, execution, auditing, payment — flows from that file.

Charter → CEO (plan) → CFO (approve budget) → Workers (execute) → Auditor (verify) → Ledger (record)

The Ledger is append-only. The Auditor is cryptographically bound. Neither can be overridden at runtime.


Why Sovereign-OS?

Typical agent frameworks Sovereign-OS
Cost control API key = burn until empty. Every cent and token is tracked. CFO approves before every task. Daily caps enforced.
Quality Hope the output is correct. Every task verified against Charter KPIs. Fail audit → TrustScore drops. Proof hash on every report.
Permissions All-or-nothing. Agents start sandboxed. They earn capabilities (SPEND_USD, CALL_API, WRITE_FILES) via TrustScore.
Monetization None. Built-in job queue, Stripe integration, ingest from any source, webhook on delivery.

Quick Start

CLI — run a single mission in 30 seconds:

git clone https://github.com/Justin0504/Sovereign-OS.git && cd Sovereign-OS
pip install -e .
sovereign run --charter charter.example.yaml "Summarize the current AI agent landscape."

Output: task plan → CFO approval → execution → audit report → ledger entry.

Web dashboard — accept paid jobs, run 24/7:

pip install -e ".[llm]"
cp .env.example .env
# Edit .env: set OPENAI_API_KEY or ANTHROPIC_API_KEY, and optionally STRIPE_API_KEY
python -m sovereign_os.web.app
# Open http://localhost:8000

From the dashboard you can submit missions, inspect the job queue, approve or retry jobs, and monitor token usage and the audit trail in real time.


Architecture

Five layers. Data flows down; accountability flows back up.

Charter (YAML)
    └─ Governance Engine
          ├─ CEO (Strategist)    — decomposes goals into ordered tasks; prunes to budget & re-plans on failure
          ├─ CFO (Treasury)      — approves/rejects per-task spend; daily cap, runway, margin + session circuit breaker
          └─ Worker Registry     — routes each task to the right worker by skill
                └─ Workers       — execute; emit TaskResult with token usage
                      └─ Auditor — verifies output against Charter KPIs (category-tuned rubric); signed AuditReport
                            └─ Ledger (UnifiedLedger) — append-only USD + token accounting

System overview

Mermaid diagram
flowchart LR
  Charter --> CEO[CEO · Plan]
  CEO --> CFO[CFO · Approve]
  CFO --> Registry[Workers]
  Registry --> Ledger[Ledger]
  Registry --> Auditor[Auditor]
  Auditor --> CFO
Loading

Key design decisions:

  • The Charter is the single source of truth. Changing runtime behavior means changing the Charter, not the code.
  • The Ledger is append-only JSONL. No delete, no update. Every entry carries a sequence number.
  • TrustScore gates capability grants. A new agent cannot spend USD or write files until it earns them through passing audits. High-risk capabilities are further scoped by just-in-time leases — granted per task, auto-expiring, revoked when the task ends (zero standing privilege).
  • Workers are pluggable. The 16 built-in workers cover common content and code tasks. You can add your own in a single Python file.
Charter panel

Marketplace oversight & delivery categories

Beyond running its own missions, Sovereign-OS acts as a governance layer over the agent-task economy — in both directions, applying the same two gates (budget + delivery quality) wherever work flows.

INBOUND   human/agent posts ─▶ ingest (TaskBounty · StacksTasker · BotBounty · ClawTasks · APB/x402)
          ─▶ [CFO budget] ─▶ governed workers ─▶ [Auditor quality] ─▶ submit

OUTBOUND  Sovereign-OS posts ─▶ [CFO budget gate] ─▶ fund escrow (RentAHuman)
          ─▶ worker delivers ─▶ [Auditor quality gate] ─▶ release | dispute
  • Inbound sources (ingest_bridge/) pull open tasks from real marketplaces (live-validated APIs); Sovereign-OS governs its own agents doing the work. Includes an APB (Agent Payment Bounty) source that discovers x402-native bounties from publishers' /.well-known/bounties.json — the standard the x402 Foundation rail is built on (see docs/AUTONOMY.md).
  • Outbound oversight (oversight/) posts tasks for external workers: the CFO approves the spend and reserves escrow before funding, and the Auditor must pass the deliverable before payment is released (else it's disputed/refunded). All money-moving calls are dry-run until explicitly set live; run python -m sovereign_os.oversight.rentahuman_preflight for a go-live check.

A task-category backbone (agents/categories.py) ties delivery together: each category (coding, data, design, email, research, writing, automation) maps to a top-tier worker, a risk-scaled budget ceiling (CategoryBudgetPolicy), a permission tier earned per category (per-category TrustScore), and the connectors it needs (connectors/, MCP / built-in / HTTP, with readiness). Inspect it with sovereign categories and sovereign connectors.

See docs/OVERSIGHT.md, docs/CATEGORIES.md, docs/CONNECTORS.md, docs/CLAWTASKS.md, docs/COST.md, and docs/X402.md.

Capability map

 INBOUND platforms                     GOVERNANCE                         AGENTIC WORKERS (by category)
 TaskBounty · StacksTasker  ──ingest──▶ ┌───────────────┐  ──route──▶  coding   → read·write·run_tests·submit_PR
 BotBounty · ClawTasks                  │ CFO budget gate│              research → web_fetch
                                        │  (per-category │              data     → web_fetch
 OUTBOUND (agent hires)                 │   ceilings)    │              design   → read_figma · generate_image
 RentAHuman escrow  ◀──post/fund──────  │ Auditor quality│              writing  → draft → self-critique → revise
   release | dispute  ◀──gate───────────│  (rubric)      │              ─────────────────────────────────────
                                        └───────────────┘              tools dispatched via connectors/, code
 Ledger (append-only, thread-safe) · per-model/category cost · TrustScore earned per category · Docker sandbox

Every task is categorized → routed to a top-tier worker → budget-gated by a per-category ceiling → executed with real tools (web/repo/Figma/image, in a Docker sandbox for untrusted code) → quality-gated before delivery or payout. All money/exec is dry-run by default.

Try it (no keys, no funds):

python examples/coding_bounty_demo.py   # bug-fix: route → budget → read·write·test·PR → audit
python examples/oversight_demo.py       # outbound: budget gate + quality gate (release vs dispute)
python examples/category_demo.py        # category → worker · budget · permission · connectors
python examples/oversight_e2e.py        # full readiness check (inbound + outbound + preflights)
sovereign categories                    # the delivery-category table
sovereign connectors                    # connector readiness + required MCP servers

Features

Governance

  • Charter-driven: mission, competencies, KPIs, fiscal boundaries — all in one YAML.
  • CEO (Strategist) decomposes natural-language goals into executable task plans with dependencies. Reactive: prune_to_budget sheds lowest-value tasks to fit a budget without breaking the dependency DAG; corrective_task builds a high-priority retry carrying the audit's failure reason and fix.
  • Profitability-first task selection — before spending any compute, an ingest screen (governance/economics.py) estimates a task's fully-loaded cost (LLM by category × complexity + settlement fee + gas) and skips anything that can't clear the margin floor. Stops the classic autonomous-agent failure mode where fees/gas turn shipped work into a net loss. Opt in with SOVEREIGN_PROFIT_SCREEN=true; tune with SOVEREIGN_{SETTLEMENT_FEE_RATIO,GAS_COST_CENTS,MIN_MARGIN_RATIO}.
  • Reactive self-repair — a task that fails audit is automatically re-run with the Auditor's reason + fix folded into its brief (SOVEREIGN_MAX_REPAIR_ATTEMPTS=N), recovering quality with no human in the loop. Only a fully-passing mission reaches delivery/settlement — a failed audit never gets submitted to a platform or charged.
  • CFO (Treasury) enforces max_task_cost_usd, daily_budget_usd, runway_days, and min_job_margin_ratio before any task runs. A runtime SpendCircuitBreaker adds fast-fail during a session — it halts the loop when cumulative spend hits a session ceiling, audits fail past a streak limit, or ROI collapses (the guard that stops runaway agent loops the pre-flight gates miss). Configure via SOVEREIGN_SESSION_CEILING_CENTS, SOVEREIGN_MAX_CONSECUTIVE_FAILURES, SOVEREIGN_ROI_FLOOR (all default off); watch it live on the dashboard Guardrails tab (session spend vs. ceiling, failure streak, ROI, one-click reset).
  • Permissions — TrustScore-gated capabilities (READ_FILES, WRITE_FILES, EXECUTE_SHELL, SPEND_USD, CALL_EXTERNAL_API), earned per category, with graduated autonomous-spend ceilings. High-risk grants use just-in-time leases: scoped to one task, bounded by TTL and use-count, auto-revoked on completion (grant_lease / use_lease / revoke_task_leases). Active leases and per-agent trust are shown live on the dashboard Guardrails tab.
  • Two dashboards, same guardrails — the web dashboard (Guardrails tab + per-task quality scorecard) and a Textual terminal Command Center (python -m sovereign_os.ui.app) both surface the CFO circuit breaker (session spend vs. ceiling, failure streak, ROI; press b to reset), active JIT leases, and the category-rubric breakdown streamed live under each audit verdict.
  • Prometheus metrics — all three guardrails export on /metrics for Grafana: breaker session spend/ceiling/ROI/trips, active JIT leases, per-agent trust, and per-category audit-quality histograms (overall + per rubric criterion). See docs/METRICS.md for the metric list and PromQL starter panels.

Execution

  • 16 built-in workers: summarize, research, reply, write_article, write_email, write_post, meeting_minutes, translate, rewrite_polish, collect_info, extract_structured, spec_writer, solve_problem, assistant_chat, code_assistant, code_review.
  • Verification-driven coding — the coding worker runs a harness-enforced loop (run_with_verified_tools): it won't accept "done" until the test suite actually passes. A premature answer bounces back with the failing test output and the model must fix and re-verify — code that can't reach green is marked tests_verified=false / success=false, so broken work never ships to a paid bounty. Gates only when execution is enabled (SOVEREIGN_CODE_EXEC_ENABLED); otherwise it's a no-op skip.
  • Multi-model: Strategist and workers can use different backends (e.g. GPT-4o for planning, Claude for execution).
  • Pluggable agent backends — for complex delivery a worker can delegate a whole task to a purpose-built coding agent (Claude Code, OpenAI Codex, Gemini CLI, Aider, or any headless CLI via a command template) instead of a single chat call. One stable AgentBackend seam; new agents plug in by config (SOVEREIGN_AGENT_BACKEND / SOVEREIGN_BACKEND_<CATEGORY>), dry-run until SOVEREIGN_AGENT_BACKEND_ENABLED=true.
  • MCP connectors in every worker — register MCP servers via SOVEREIGN_MCP_SERVERS and their tools appear inline in every worker's tool-use loop, next to built-ins like web_fetch. Any MCP-provided connector (DBs, search, SaaS, another agent over MCP) becomes callable mid-task with no code change — the universal connector layer both Claude Agent SDK and Codex speak. See docs/BACKENDS.md.
  • Dynamic worker loading: drop a Python file into sovereign_os/agents/user_workers/ — no registration boilerplate.

Auditing

  • Every task output is verified against Charter KPIs by the ReviewEngine.
  • Category-tuned analytic rubric: high-value deliverables are scored criterion-by-criterion on a rubric matched to the work type — coding gets correctness/completeness/robustness/relevance, writing gets clarity/voice, etc. — always ending in a safety criterion. Criterion order is deterministically shuffled per task to blunt LLM-judge positional bias.
  • Value-aware bar: higher-paid jobs must clear a stricter passing score; cheap tasks can skip the (relatively expensive) LLM judge.
  • AuditReport carries score, passed, reason, suggested_fix, sub_scores, and proof_hash (SHA-256 of inputs + output). Each job's delivery view renders a quality scorecard — the per-criterion rubric breakdown (bars per criterion, category label, overall score) so you can see why a deliverable passed or failed, not just the verdict.
  • Append-only audit trail (JSONL). Integrity verifiable offline.

Monetization & job queue

  • SQLite-backed job queue (Redis optional for multi-instance).
  • Stripe integration: set STRIPE_API_KEY and each completed job triggers a real charge; transactions appear in your Stripe Dashboard.
  • Auto-approval mode or manual review per job.
  • Ingest from any HTTP endpoint, Reddit (PRAW), Shopify, WooCommerce, or custom scrapers via the ingest bridge.
  • Webhook delivery: POST job result to any URL on completion.

Job delivery result  Stripe Dashboard — completed charges

Left: job result delivered in the dashboard. Right: charges recorded directly in Stripe Dashboard.

Observability

  • OpenTelemetry tracing.
  • Prometheus metrics at GET /metrics: job counters, queue depth, task duration histograms.
  • GET /health returns config warnings (missing API keys, payment mode).
  • Structured JSON logs with correlation IDs.

Security

  • API key authentication (constant-time comparison).
  • Optional IP allowlist and per-IP rate limiting.
  • Job input validation (Pydantic v2).
  • No secrets in code; everything via environment variables.

TrustScore gates  Permission gate detail

TrustScore-gated permission system — agents earn capabilities through passing audits.


Project Layout

sovereign_os/
├── models/           # Charter schema (Pydantic v2)
├── ledger/           # UnifiedLedger — append-only USD + token accounting
├── governance/       # CEO (Strategist), CFO (Treasury), SpendCircuitBreaker, GovernanceEngine
├── agents/           # Workers, WorkerRegistry, SovereignAuth (+ JIT leases), categories, user_workers/
├── auditor/          # ReviewEngine, category rubric, AuditReport (proof_hash), KPIValidator
├── jobs/             # JobStore (SQLite), RedisJobStore
├── ingest/           # HTTP poller → job queue
├── ingest_bridge/    # Reddit, scrapers, Shopify/WooCommerce → jobs
├── oversight/        # Outbound task posting + escrow (RentAHuman, StacksTasker)
├── delivery/         # Deliver-back adapters (TaskBounty, StacksTasker, ClawTasks, Reddit, APB/x402)
├── connectors/       # web_fetch, code_workspace, figma, image_gen, submit_pr, sandbox
├── payments/         # StripePaymentService, X402PaymentService, DummyPaymentService
├── ui/               # Textual terminal Command Center (TaskTree · DecisionStream · Finance & Guardrails)
├── web/              # FastAPI app, dashboard, /api/jobs, /health, Stripe webhook
└── telemetry/        # OpenTelemetry, Prometheus
charters/             # Example Charter YAML files
examples/             # Example job payloads
tests/                # pytest
docs/                 # Full documentation

Deployment

Local (bare metal)

git clone https://github.com/Justin0504/Sovereign-OS.git
cd Sovereign-OS
pip install -e ".[llm]"
cp .env.example .env
# Edit .env with your API keys
python -m sovereign_os.web.app

Listens on http://0.0.0.0:8000 by default. Set SOVEREIGN_HOST and SOVEREIGN_PORT to change.

Docker Compose (recommended for production)

cp .env.example .env
# Edit .env
docker compose up -d
# Web UI:   http://localhost:8000
# Metrics:  http://localhost:8000/metrics
# Health:   http://localhost:8000/health

docker-compose.yml includes the web app, Redis (for multi-instance job queue), and named volumes for the ledger and job database.

# docker-compose.yml excerpt
services:
  web:
    build: .
    ports: ["8000:8000"]
    env_file: .env
    volumes:
      - sovereign_data:/app/data
  redis:
    image: redis:7-alpine
    volumes:
      - redis_data:/data

See docs/DEPLOY.md for volume strategy, health checks, and graceful shutdown.

Ingest bridge (optional)

Pull real orders from Reddit, scrapers, or Shopify:

pip install -e ".[bridge]"
python -m sovereign_os.ingest_bridge   # serves on :9000
# In .env: SOVEREIGN_INGEST_URL=http://localhost:9000/jobs?take=true

See docs/INGEST_BRIDGE.md for Reddit credentials, scraper targets, and retail connectors.


Configuration

All configuration is via environment variables. Copy .env.example to .env.

Variable Required Description
OPENAI_API_KEY One of these OpenAI API key for LLM workers
ANTHROPIC_API_KEY One of these Anthropic API key for LLM workers
STRIPE_API_KEY Optional Stripe secret key (sk_test_… or sk_live_…). Without this, charges are simulated.
SOVEREIGN_CHARTER_PATH Optional Path to Charter YAML. Default: charter.default.yaml.
SOVEREIGN_JOB_DB Optional SQLite path. Default: data/jobs.db.
SOVEREIGN_AUTO_APPROVE_JOBS Optional true to skip manual approval. Default: false.
SOVEREIGN_JOB_WORKER_ENABLED Optional true to start the background job processor. Default: false.
SOVEREIGN_INGEST_URL Optional HTTP endpoint to poll for incoming jobs.
SOVEREIGN_INGEST_ENABLED Optional true to enable the ingest poller. Default: false.
SOVEREIGN_API_KEY Optional Bearer token for /api/* endpoints.
SOVEREIGN_ALLOWED_IPS Optional Comma-separated IP allowlist.
SOVEREIGN_WEBHOOK_URL Optional URL to POST job results to on completion.
REDIS_URL Optional Redis connection string for multi-instance job queue.

Full reference: docs/CONFIG.md.


Custom Workers

Workers are the execution layer. Each worker maps to a skill name and receives a task description; it returns a TaskResult with output and token usage.

# sovereign_os/agents/user_workers/my_worker.py
from sovereign_os.agents.base import BaseWorker, TaskResult

class MyWorker(BaseWorker):
    skill = "my_skill"

    async def run(self, task_description: str, **kwargs) -> TaskResult:
        result = await self._chat(f"Do this: {task_description}")
        return TaskResult(output=result, success=True)

Drop the file into sovereign_os/agents/user_workers/ — it is loaded automatically on startup. Reference it in your Charter:

competencies:
  - name: my_skill
    description: "Does the custom thing."

See docs/WORKER.md for the full API, LLM helper methods, and advanced patterns.


Testing

pip install -e ".[dev]"
pytest tests/ -v

Tests (257, all passing) cover the governance engine, CFO circuit breaker, JIT capability leases, category-tuned audit rubric, reactive planning, ledger accounting, marketplace oversight, payments (Stripe + x402), job store, compliance hooks, web API, and webhook delivery. Run with LLM keys unset so tests use stubs (no network).


Docs

Document Description
QUICKSTART.md Stripe + LLM key, first paid job, curl examples, troubleshooting.
CONFIG.md All SOVEREIGN_* environment variables.
CHARTER.md How to write mission, competencies, KPIs, and fiscal bounds.
WORKER.md How to write and register custom workers.
DEPLOY.md Docker Compose, volumes, health checks, graceful shutdown.
INGEST_BRIDGE.md Pull orders from Reddit, scrapers, Shopify/WooCommerce.
MONETIZATION.md Job queue, Stripe charging, approval flow, compliance.
POSITIONING.md Order sources (open platforms, partners, API), service model.
CEO_CFO_PROFITABILITY.md Unit economics: margin floor, rejecting unprofitable jobs.
ARCHITECTURE_FAQ.md Design decisions and trade-offs.
AUDIT_PROOF.md Verifiable audit trail, proof_hash, integrity check.
METRICS.md Prometheus metrics (breaker, JIT leases, audit quality) + Grafana/PromQL starter.
BACKENDS.md Pluggable agent backends (Claude Code / Codex / CLI agents), MCP compatibility, adapting to platform updates.
AUTONOMY.md Human-out-of-loop runbook: profit screening, self-repair, platform notes, safety switches.
MULTI_INSTANCE.md Redis queue, horizontal scaling, concurrency.
BACKUP.md Back up job DB and ledger for disaster recovery.
FUTURE_DEVELOPMENT.md Roadmap: stability, UX, workers, scale, community.
CONTRIBUTING.md How to contribute, run tests, submit PRs.
SECURITY.md Vulnerability disclosure policy.

Contributing

Pull requests are welcome. Run the test suite before submitting:

pip install -e ".[dev]"
pytest tests/ -v

See CONTRIBUTING.md for the full process.


License

MIT — see LICENSE.


Sovereign-OS — Think. Approve. Execute. Audit. Every time.

About

Constitution-first AI orchestration: one Charter (YAML) defines mission, budget & rules. CEO plans → CFO approves → Ledger tracks every cent & token → Auditor scores. 16 workers, Stripe, MCP. Think. Audit. Execute.

Topics

Resources

Contributing

Security policy

Stars

92 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages