← All changelog entries

Model scenarios can be saved as repeatable benchmarks

Labs can preserve a named model-and-template scenario, run it again, and compare its result with earlier runs.

Labs scenarios can now be saved as named benchmarks and run again with one action. Each run resolves the benchmark’s templates, enabled models, and current live parser version, then passes through the same estimate and budget gate as an ad-hoc scenario.

Benchmark cards show the latest per-model results and an oldest-to-newest trend across completed runs. Retiring a saved benchmark keeps its run history navigable.