awesome-agent-benchmarks
๐ง Discover and evaluate advanced benchmark datasets for Large Language Model agents to enhance performance assessment in real-world tasks.
Why this rank:Recent releaseHealthy release cadenceStrong adoption
Description
๐ง Discover and evaluate advanced benchmark datasets for Large Language Model agents to enhance performance assessment in real-world tasks.
Release History
| Version | Changes | Urgency | Date |
|---|---|---|---|
| master@2026-08-30 | Latest activity on master branch | High | 8/30/2026 |
| 0.0.0 | No release found โ using repo HEAD | High | 4/9/2026 |
Dependencies & License Audit
Loading dependencies...
Similar Packages
opentulpaSelf-hosted personal AI agent that lives in your DMs. Describe any workflow: triage Gmail, pull a Giphy feed, build a Slack bot, monitor markets. It writes the code, runs it, schedules it, and saves imain@2026-09-03
CopilotKitThe Frontend Stack for Agents & Generative UI. React + Angular. Makers of the AG-UI Protocolv1.71.0
claude-code-tipsProvide ready-to-use plugins, hooks, and commands to enhance Claude Code sessions with data mining, automation, and integration tools.main@2026-09-09
octobenchBenchmark and compare LLM tool, configuration, and prompt setups using a shared case framework with automated scoring and telemetry.main@2026-09-09
agentic-rag๐ Enable smart document and data search with AI-powered chat, vector search, and SQL querying across multiple file formats.main@2026-09-09
More in Testing
ContribAIAutonomous AI agent that contributes to open source โ discovers repos, analyzes code, generates fixes, and submits PRs
ObservalObserval is an AI agent registry with first in class observabilty and eval framework
GitoAn AI-powered GitHub code review tool that uses LLMs to detect high-confidence, high-impact issuesโsuch as security vulnerabilities, bugs, and maintainability concerns.
trulensEvaluation and Tracking for LLM Experiments and AI Agents
