EVALUATION LABDeterministic benchmark · v1
Back to workspace

EVALUATION

Can the agent reach the right business conclusion?

30 labelled business questions test intent, metric selection and root-cause attribution.

30Evaluation queries
LockedGround truth
Benchmark v1Version
Ready0 / 30

BENCHMARK RESULTS

Measured performance

Not run
Run the benchmark to generate measured accuracy and failure cases.

ABOUT THE BENCHMARK

30 labelled business questionsSynthetic evaluation datasetLocked ground truthDeterministic data functions