Skip to content
#

ai-agent-testing

Here are 9 public repositories matching this topic...

Language: All
Filter by language

First-Person Agent Memory Bench. 10 Categories including fact recall, multi-hop links, temporal reasoning, fact overwrites, speaker traps, refusal, credibility, and agentic tool usage. 540K token / 60 session corpus, all in first person. Dynamic output-answer-key portion. Comprehensive report with visuals and miss breakdown.

  • Updated Aug 27, 2026
  • Python

Production-readiness auditor for AI agents — a Claude Code skill that grades any agent repo across evaluation, observability, data foundation, orchestration & governance, then tells you exactly what to build next

  • Updated Jul 6, 2026
  • Python

Improve this page

Add a description, image, and links to the ai-agent-testing topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the ai-agent-testing topic, visit your repo's landing page and select "manage topics."

Learn more