Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
69 changes: 33 additions & 36 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![PRs Welcome](https://img.shields.io/badge/PRs-welcome-brightgreen.svg)](http://makeapullrequest.com)
[![Evaluations](https://img.shields.io/badge/Evaluations-106-blue.svg)](#-current-coverage)
[![Evaluations](https://img.shields.io/badge/Evaluations-156-blue.svg)](#-current-coverage)
[![GitHub Stars](https://img.shields.io/github/stars/Guard0-Security/TrustVector?style=social)](https://github.com/Guard0-Security/TrustVector)

TrustVector is an evidence-based evaluation framework for AI systems, providing transparent, multi-dimensional trust scores across **security**, **privacy**, **performance**, **trust**, and **operational excellence**.
Expand Down Expand Up @@ -87,57 +87,54 @@ const customScore = calculateCustomScore(claudeSonnet, {

## 📊 Current Coverage

**106 Total Evaluations** across 3 categories:
**156 Total Evaluations** across 3 categories (last refreshed June 2026):

### AI Models (38)
### AI Models (60)

**Frontier Models:**
- ✅ Claude Sonnet 4.5, Claude Opus 4.1, Claude 3.7 Sonnet, Claude 3.5 Haiku (Anthropic)
- ✅ GPT-5, GPT-4.5, GPT-4.1, GPT-4o, GPT-4o Mini (OpenAI)
- ✅ o1, o1 Mini, o3, o3 Mini (OpenAI Reasoning)
- ✅ Gemini 2.5 Pro, Gemini 2.0 Flash (Google)
- ✅ Llama 4 Behemoth, Llama 4 Maverick, Llama 4 Scout, Llama 3.3 70B, Llama 3.1 405B (Meta)
- ✅ Grok 3 Beta (xAI)
- ✅ DeepSeek R1, DeepSeek V3 (DeepSeek)

**Specialized & Open Source:**
- ✅ Gemma 3 27B (Google)
- ✅ Qwen2.5-VL 32B (Alibaba)
- ✅ Nemotron Ultra 253B (NVIDIA)
- ✅ Nova Pro (Amazon)
- ✅ Claude Fable 5, Claude Opus 4.8 / 4.7 / 4.6 / 4.5, Claude Sonnet 4.6 / 4.5, Claude Haiku 4.5 (Anthropic)
- ✅ GPT-5.5, GPT-5.4, GPT-5.3-Codex, GPT-5.2, GPT-5.1, GPT-5, o-series (OpenAI)
- ✅ Gemini 3.1 Pro, Gemini 3.5 Flash, Gemini 3 Pro/Flash (Google)
- ✅ Grok 4.3, Grok 4.1 (xAI)
- ✅ Nova 2 Lite, Nova Pro (Amazon)

**Open-Weight Models:**
- ✅ DeepSeek V4, DeepSeek V3.2, DeepSeek R1 (DeepSeek)
- ✅ Qwen3.5 (Alibaba), Kimi K2.6 (Moonshot), GLM-5 (Z.ai), MiniMax-M2
- ✅ Mistral Large 3 (Mistral), Command A+ (Cohere)
- ✅ Gemma 4, Gemma 3 (Google), gpt-oss-120b/20b (OpenAI)
- ✅ Llama 4 Maverick/Scout, Llama 3.x (Meta), Nemotron (NVIDIA)

**[See all models →](/data/models)**

### AI Agents (34)
### AI Agents (50)

**Enterprise Platforms:**
- ✅ Amazon Bedrock Agents, Azure Bot Service, Google Agent Builder
- ✅ IBM Watson Assistant, Google Dialogflow, Amazon Lex
**Coding & Autonomous Agents:**
- ✅ Claude Code + Claude Agent SDK (Anthropic), OpenAI Codex, Devin (Cognition)
- ✅ Cursor, GitHub Copilot coding agent, Google Jules, Gemini CLI, Manus

**Developer Frameworks:**
- ✅ LangGraph Agent, LlamaIndex Agent, CrewAI, AutoGen
- ✅ Haystack, LangFlow, Flowise, E2B Agents
- ✅ OpenAI Agents SDK, Google ADK, Microsoft Agent Framework, AWS Strands Agents
- ✅ LangGraph, CrewAI, LlamaIndex, Pydantic AI, smolagents, Mastra, Dify

**Autonomous Agents:**
- ✅ AutoGPT, BabyAGI, AgentGPT, Adala
- ✅ And 15+ more...
**Enterprise Platforms:**
- ✅ Amazon Bedrock Agents, Azure Bot Service, Gemini Enterprise Agent Platform
- ✅ IBM watsonx Assistant, Google Dialogflow, Amazon Lex, and more

**[See all agents →](/data/agents)**

### MCP Servers (34)

**Cloud & Infrastructure:**
- ✅ AWS, Azure, Cloudflare, Docker, Kubernetes
### MCP Servers (46)

**Development Tools:**
- ✅ GitHub, Git, Filesystem, Memory
**Top Ecosystem Servers:**
- ✅ Context7, Chrome DevTools MCP, Playwright MCP, Serena

**Productivity & Business:**
- ✅ Gmail, Google Drive, Calendar, Linear, Atlassian
- ✅ Datadog, Elasticsearch, MongoDB
**Official Vendor Servers:**
- ✅ GitHub, Figma, Stripe, Notion, Vercel, Hugging Face, Zapier, Apify

**Utilities:**
- ✅ Brave Search, Fetch, Everything
**Reference & Community:**
- ✅ Fetch, Git, Filesystem, Memory, Sequential Thinking, Time, Everything
- ✅ AWS, Azure, Cloudflare, Docker, Kubernetes, databases, and more
- ⚠️ Archived reference servers (Puppeteer, Postgres, SQLite, Slack, …) are flagged with security advisories
- ✅ And 15+ more...

**[See all MCPs →](/data/mcps)**
Expand Down
40 changes: 28 additions & 12 deletions data/agents/agentgpt.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@
"name": "AgentGPT",
"provider": "Reworkd",
"version": "Platform",
"last_evaluated": "2025-11-09",
"last_evaluated": "2026-06-10",
"evaluated_by": "TrustVector Team",
"description": "Browser-based autonomous AI agent platform for deploying and managing GPT-powered agents. Enables users to create goal-oriented autonomous agents that break down objectives and execute tasks without continuous human intervention.",
"description": "DISCONTINUED: the AgentGPT repository was archived on 2026-01-28 (last release v1.0.0, Nov 2023) and the hosted site is frozen. Formerly a browser-based autonomous AI agent platform that let users create goal-oriented agents which break down objectives and execute tasks without continuous human intervention. Not recommended for new use.",
"website": "https://agentgpt.reworkd.ai/",
"trust_vector": {
"performance_reliability": {
Expand Down Expand Up @@ -249,7 +249,7 @@
}
},
"trust_transparency": {
"overall_score": 79,
"overall_score": 76,
"criteria": {
"documentation_quality": {
"score": 75,
Expand Down Expand Up @@ -308,23 +308,30 @@
"last_verified": "2025-11-09"
},
"community_support": {
"score": 73,
"confidence": "medium",
"score": 55,
"confidence": "high",
"evidence": [
{
"source": "Community",
"url": "https://github.com/reworkd/AgentGPT/discussions",
"date": "2024-10-15",
"value": "Active GitHub community and Discord"
},
{
"source": "GitHub Repository Status",
"url": "https://github.com/reworkd/AgentGPT",
"date": "2026-06-10",
"value": "Repository archived 2026-01-28; community activity has ceased"
}
],
"methodology": "Community engagement analysis",
"last_verified": "2025-11-09"
"last_verified": "2026-06-10",
"notes": "Score reduced: repository archived, no further community development"
}
}
},
"operational_excellence": {
"overall_score": 70,
"overall_score": 65,
"criteria": {
"ease_of_use": {
"score": 88,
Expand Down Expand Up @@ -383,18 +390,25 @@
"last_verified": "2025-11-09"
},
"production_readiness": {
"score": 60,
"confidence": "medium",
"score": 40,
"confidence": "high",
"evidence": [
{
"source": "Platform Maturity",
"url": "https://github.com/reworkd/AgentGPT",
"date": "2024-10-01",
"value": "Experimental platform, not designed for production use"
},
{
"source": "GitHub Repository Status",
"url": "https://github.com/reworkd/AgentGPT",
"date": "2026-06-10",
"value": "Repository archived 2026-01-28; last release v1.0.0 (Nov 2023); hosted site frozen"
}
],
"methodology": "Production readiness assessment",
"last_verified": "2025-11-09"
"last_verified": "2026-06-10",
"notes": "Score reduced: project discontinued and repository archived on 2026-01-28"
},
"reliability": {
"score": 62,
Expand Down Expand Up @@ -474,7 +488,8 @@
"Limited tool access and action capabilities",
"Can be expensive with unpredictable OpenAI API costs",
"No enterprise features (auth, monitoring, management)",
"Reliability and success rate varies significantly"
"Reliability and success rate varies significantly",
"Discontinued: repository archived 2026-01-28, hosted site frozen, no further updates"
],
"metadata": {
"license": "GPL-3.0",
Expand All @@ -500,6 +515,7 @@
},
"tags": [
"autonomous",
"web-based"
"web-based",
"archived"
]
}
18 changes: 14 additions & 4 deletions data/agents/amazon-bedrock-agents.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@
"name": "Amazon Bedrock Agents",
"provider": "Amazon Web Services",
"version": "2024",
"last_evaluated": "2025-11-09",
"last_evaluated": "2026-06-10",
"evaluated_by": "TrustVector Team",
"description": "Fully managed AWS service for building and deploying generative AI agents. Handles orchestration, memory, knowledge bases, and action groups with enterprise security and scalability built-in.",
"description": "Fully managed AWS service for building and deploying generative AI agents with orchestration, memory, knowledge bases, and action groups. Note: AWS's strategic agent runtime is now Bedrock AgentCore (GA 2025-10-13; framework-agnostic, 8-hour sessions, session isolation) paired with the open-source Strands Agents SDK; evaluate AgentCore for new builds.",
"website": "https://aws.amazon.com/bedrock/agents/",
"trust_vector": {
"performance_reliability": {
Expand Down Expand Up @@ -378,10 +378,16 @@
"url": "https://aws.amazon.com/bedrock/customers/",
"date": "2024-10-01",
"value": "Production-ready managed service with enterprise customers"
},
{
"source": "Amazon Bedrock AgentCore GA Announcement",
"url": "https://aws.amazon.com/about-aws/whats-new/2025/10/amazon-bedrock-agentcore-available/",
"date": "2026-06-10",
"value": "AWS's strategic agent runtime is now Bedrock AgentCore (GA 2025-10-13): framework-agnostic, 8-hour sessions, session isolation; complemented by the open-source Strands Agents SDK"
}
],
"methodology": "Production readiness assessment",
"last_verified": "2025-11-09"
"last_verified": "2026-06-10"
}
}
}
Expand Down Expand Up @@ -448,7 +454,8 @@
"Limited to AWS Bedrock foundation models",
"Less flexibility than code-based frameworks",
"Requires AWS expertise for optimal configuration",
"Not open source, limited customization of core orchestration"
"Not open source, limited customization of core orchestration",
"AWS's strategic focus has shifted to Bedrock AgentCore (GA 2025-10-13) and the Strands Agents SDK; new agent workloads should evaluate AgentCore first"
],
"metadata": {
"license": "Proprietary (AWS)",
Expand All @@ -471,6 +478,9 @@
"regions_available": "Multiple AWS regions",
"sla": "99.9%"
},
"related_entities": [
"strands-agents"
],
"tags": [
"aws"
]
Expand Down
21 changes: 16 additions & 5 deletions data/agents/autogen.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@
"name": "Microsoft AutoGen",
"provider": "Microsoft Research",
"version": "0.4",
"last_evaluated": "2025-11-09",
"last_evaluated": "2026-06-10",
"evaluated_by": "TrustVector Team",
"description": "Multi-agent conversation framework enabling next-gen LLM applications with conversable agents that can operate in various modes combining LLMs, human inputs, and tools. Supports complex workflows through agent conversations.",
"description": "MAINTENANCE MODE: AutoGen now receives bug/security fixes only and is superseded by the Microsoft Agent Framework (1.0 GA on 2026-04-03), the recommended migration path. AutoGen is a multi-agent conversation framework for LLM applications with conversable agents combining LLMs, human input, and tools across complex workflows.",
"website": "https://microsoft.github.io/autogen/",
"trust_vector": {
"performance_reliability": {
Expand Down Expand Up @@ -377,10 +377,16 @@
"url": "https://github.com/microsoft/autogen",
"date": "2024-10-20",
"value": "Very active community with Microsoft backing"
},
{
"source": "Microsoft Agent Framework Migration Guidance",
"url": "https://devblogs.microsoft.com/semantic-kernel/migrate-your-semantic-kernel-and-autogen-projects-to-microsoft-agent-framework-release-candidate/",
"date": "2026-06-10",
"value": "AutoGen is in maintenance mode (bug/security fixes only); Microsoft Agent Framework 1.0 reached GA on 2026-04-03 as the successor"
}
],
"methodology": "Community activity analysis",
"last_verified": "2025-11-09"
"last_verified": "2026-06-10"
}
}
}
Expand Down Expand Up @@ -447,7 +453,8 @@
"Requires careful prompt engineering for agent roles",
"Limited built-in persistence for long-running workflows",
"Some learning curve for advanced features",
"Performance depends heavily on LLM quality"
"Performance depends heavily on LLM quality",
"Maintenance mode: bug/security fixes only; new development targets Microsoft Agent Framework"
],
"metadata": {
"license": "Apache 2.0",
Expand All @@ -474,9 +481,13 @@
"contributors": "559+",
"transition_notice": "Microsoft Agent Framework is the recommended path forward; AutoGen receives maintenance and critical patches only"
},
"related_entities": [
"microsoft-agent-framework"
],
"tags": [
"multi-agent",
"microsoft",
"open-source"
"open-source",
"maintenance-mode"
]
}
23 changes: 16 additions & 7 deletions data/agents/babyagi.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@
"name": "BabyAGI",
"provider": "Yohei Nakajima",
"version": "Classic",
"last_evaluated": "2025-11-09",
"last_evaluated": "2026-06-10",
"evaluated_by": "TrustVector Team",
"description": "Minimalist autonomous task-driven AI agent that creates, prioritizes, and executes tasks based on results of previous tasks and a predefined objective. Demonstrates AGI concepts in under 200 lines of code.",
"description": "ARCHIVED: the original BabyAGI repo was archived to babyagi_archive in September 2024 and replaced by an experimental self-building framework; it is not production-maintained. Originally a minimalist autonomous task-driven AI agent that created, prioritized, and executed tasks toward an objective, demonstrating AGI concepts in under 200 lines of code.",
"website": "https://github.com/yoheinakajima/babyagi",
"trust_vector": {
"performance_reliability": {
Expand Down Expand Up @@ -310,7 +310,7 @@
}
},
"operational_excellence": {
"overall_score": 61,
"overall_score": 57,
"criteria": {
"ease_of_integration": {
"score": 75,
Expand Down Expand Up @@ -370,18 +370,25 @@
"last_verified": "2025-11-09"
},
"production_readiness": {
"score": 50,
"score": 35,
"confidence": "high",
"evidence": [
{
"source": "Project Purpose",
"url": "https://github.com/yoheinakajima/babyagi",
"date": "2024-09-01",
"value": "Designed as concept demonstration, not production system"
},
{
"source": "GitHub Repository Status",
"url": "https://github.com/yoheinakajima/babyagi",
"date": "2026-06-10",
"value": "Original repo archived to babyagi_archive (Sept 2024); replaced by an experimental self-building framework; not production-maintained"
}
],
"methodology": "Production readiness assessment",
"last_verified": "2025-11-09"
"last_verified": "2026-06-10",
"notes": "Score reduced: original project archived and unmaintained since September 2024"
}
}
}
Expand All @@ -400,7 +407,8 @@
"Can generate excessive tasks leading to high costs",
"No built-in security or sandboxing features",
"Limited tool integration in classic version",
"Unpredictable behavior and task completion quality"
"Unpredictable behavior and task completion quality",
"Archived (Sept 2024): original repo moved to babyagi_archive with no further maintenance"
],
"metadata": {
"license": "MIT",
Expand Down Expand Up @@ -471,6 +479,7 @@
"tags": [
"autonomous",
"experimental",
"open-source"
"open-source",
"archived"
]
}
Loading
Loading