IFEval
v1writing.structuredGoogle's Instruction-Following Eval — 25 verifiable-instruction types scored by rule validators (no LLM judge). 100-task subsample of the public split, held constant to enable version-over-version comparison.
Scored by ifeval over 100 tasks. Methodology: see BENCHMARKS.md in the source repo.
| Rank | Skill | Version | Accuracy | Cost / success | Median latency | Robustness |
|---|---|---|---|---|---|---|
| — | category median | — | — | — | — | — |
| No completed runs yet — skills in this category are still being scored. | ||||||