Skip to content
#

test-evaluation

Here are 2 public repositories matching this topic...

A governed world-model evidence layer for AI agents: simulate bounded scenarios, track assumptions, score prediction-vs-reality error, and produce human-reviewable execution evidence.

  • Updated May 22, 2026
  • Python

Add this topic to your repo

To associate your repository with the test-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more