freshcrate
Skin:/
Home > Testing > awesome-agent-benchmarks

awesome-agent-benchmarks

๐Ÿง  Discover and evaluate advanced benchmark datasets for Large Language Model agents to enhance performance assessment in real-world tasks.

Why this rank:Recent releaseHealthy release cadenceStrong adoption

Description

๐Ÿง  Discover and evaluate advanced benchmark datasets for Large Language Model agents to enhance performance assessment in real-world tasks.

Release History

VersionChangesUrgencyDate
master@2026-08-30Latest activity on master branchHigh8/30/2026
0.0.0No release found โ€” using repo HEADHigh4/9/2026

Dependencies & License Audit

Loading dependencies...

Similar Packages

opentulpaSelf-hosted personal AI agent that lives in your DMs. Describe any workflow: triage Gmail, pull a Giphy feed, build a Slack bot, monitor markets. It writes the code, runs it, schedules it, and saves imain@2026-09-03
CopilotKitThe Frontend Stack for Agents & Generative UI. React + Angular. Makers of the AG-UI Protocolv1.71.0
claude-code-tipsProvide ready-to-use plugins, hooks, and commands to enhance Claude Code sessions with data mining, automation, and integration tools.main@2026-09-09
octobenchBenchmark and compare LLM tool, configuration, and prompt setups using a shared case framework with automated scoring and telemetry.main@2026-09-09
agentic-rag๐Ÿ“„ Enable smart document and data search with AI-powered chat, vector search, and SQL querying across multiple file formats.main@2026-09-09

More in Testing

ContribAIAutonomous AI agent that contributes to open source โ€” discovers repos, analyzes code, generates fixes, and submits PRs
ObservalObserval is an AI agent registry with first in class observabilty and eval framework
GitoAn AI-powered GitHub code review tool that uses LLMs to detect high-confidence, high-impact issuesโ€”such as security vulnerabilities, bugs, and maintainability concerns.
trulensEvaluation and Tracking for LLM Experiments and AI Agents