Skip to content
#

ai-benchmarking

Here are 19 public repositories matching this topic...

ai-agents-reality-check

Benchmarking the gap between AI agent hype and architecture. Three agent archetypes, 73-point performance spread, stress testing, network resilience, and ensemble coordination analysis with statistical validation.

  • Updated Apr 2, 2026
  • Python

Uncertainty & Confidence Management (UCM): A healthcare AI benchmark suite for uncertainty recognition, justification boundaries, confidence calibration, proportionate action, and reassessment.

  • Updated Jul 14, 2026

🔬 Research Project: An automated framework to generate, configure, and evaluate multi-agent AI crews for financial modeling using a Meta-Agent pipeline. This study evaluates the performance of dynamically synthesized MAS (Multi-Agent Systems) against manual expert-defined benchmarks in financial risk contexts.

  • Updated May 24, 2026
  • Python

Improve this page

Add a description, image, and links to the ai-benchmarking topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the ai-benchmarking topic, visit your repo's landing page and select "manage topics."

Learn more