octobench
Benchmark and compare LLM tool, configuration, and prompt setups using a shared case framework with automated scoring and telemetry.
Why this rank:Recent releaseHealthy release cadenceStrong adoption
Description
Benchmark and compare LLM tool, configuration, and prompt setups using a shared case framework with automated scoring and telemetry.
README
Release History
| Version | Changes | Urgency | Date |
|---|---|---|---|
| main@2026-09-09 | Latest activity on main branch | High | 9/9/2026 |
| 0.0.0 | No release found โ using repo HEAD | High | 4/9/2026 |
Dependencies & License Audit
Loading dependencies...
Similar Packages
hatch3rInstall an agentic coding setup that adds multiple AI agents, skills, and rules to enhance automation across GitHub, Azure DevOps, or GitLab repositories.main@2026-09-09
simBuild, deploy, and orchestrate AI agents. Sim is the central intelligence layer for your AI workforce.v0.8.20
sdk-pythonA model-driven approach to building AI agents in just a few lines of code.typescript/v1.16.0
autonomous-agentic-research-swarmFile-based autonomous agentic research swarm template (Planner/Worker/Judge) with contracts, workstreams, and deterministic quality gates.main@2026-08-05
More in Testing
ContribAIAutonomous AI agent that contributes to open source โ discovers repos, analyzes code, generates fixes, and submits PRs
ObservalObserval is an AI agent registry with first in class observabilty and eval framework
GitoAn AI-powered GitHub code review tool that uses LLMs to detect high-confidence, high-impact issuesโsuch as security vulnerabilities, bugs, and maintainability concerns.
trulensEvaluation and Tracking for LLM Experiments and AI Agents
