Gorio Tech Blog search

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness | Summary

|

Contents

This article explains the key points of Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness.

  • 2026-07-21 (arXiv)
  • Chen, Xilun, Feizollahi, Zhaleh, Goodwin, Ross, Moon, Seungwhan, Yih, Scott, Donmez, Pinar, Damavandi, Babak, Dong, Luna.
  • Meta AI
  • Paper
  • Github

Read this article in Korean


Summary

  • GAMUT evaluates factual completeness in open-ended long-form answers, rather than only the precision of claims that an answer makes. It contains 1,813 image-grounded everyday deep-research questions from 10 domains, each paired with an expert-verified, evidence-backed rubric; among 14 evaluated models, Gemini 3.1 Pro achieves the highest score, 58.7%.

3 A Two-Level Meta-Rubric Framework for Evaluating Open-Ended Generation

The framework separates expressive rubric authoring from reliable automated grading. A structured meta-rubric records the organization of a complete answer, and fixed mechanical rules compile it into self-contained binary checks for an LLM judge.

3.1 Meta-Rubrics

Each meta-rubric item has an importance tier—Answer-Critical, Valuable, or Context—and one of five types: Simple Knowledge, Strict List, Flexible List, Process, or Relationship. Lists and processes may also include a meta-insight, a synthesis across items that cannot be reduced to an individual fact.

3.2 From Meta-Rubrics to Binary Rubrics

Simple Knowledge becomes one check, while Strict Lists become one check per required item. Flexible Lists receive a self-contained coverage threshold plus lower-tier credit for broader coverage; Processes yield step-presence and ordering checks, while Relationships and meta-insights are decomposed into atomic checks.

A Jollof rice example traces human-verified snippets into a typed meta-rubric and then into tiered binary checks. Mandatory ingredients become individual critical checks, supporting ingredients become a coverage threshold plus lower-tier item credit, and preparation becomes both step and ordering checks.

An everyday deep-research question, supporting web evidence, its structured meta-rubric, and the compiled checklist organized by importance.
An everyday deep-research question, supporting web evidence, its structured meta-rubric, and the compiled checklist organized by importance.

3.3 Scoring

For every binary check, the judge returns meets, partially meets, missing, or contradicts. Per-tier scores use λ = 0.5 for partial credit and µ = −2 for contradictions, then combine Answer-Critical, Valuable, and Context tiers with weights 0.6, 0.3, and 0.1; a missing tier’s weight is redistributed across the remaining tiers.

4 Dataset Construction

Dataset construction combines frontier-LLM proposals with multi-round expert human review. Questions are created from real wearable imagery, while rubrics are researched, structured, compiled, refined, and checked against cited web evidence before release.

4.1 Question Creation

The authors begin with 1,938 CRAG-MM images, about 80% of which are egocentric smart-glasses photographs, spanning 10 everyday domains. Gemini 3.1 Pro proposes research-intensive image-dependent questions; two annotators review each candidate, a third adjudicates disagreements, and an automatic pass filters, ranks, and diversifies accepted questions.

4.2 Rubric Creation

For each accepted question, Gemini 3.1 Pro researches the web, drafts a meta-rubric with supporting snippets, and compiles it into binary checks. A self-refinement pass re-verifies sources and independently searches for missing evidence; in-house annotators then revise rubrics and citations over two human-review rounds. Corrections to 12 examples and removal of unsound cases leave 1,813 questions.

5 Evaluating Models on Gamut

Models answer from the image and question, and Gemini 3.1 Pro judges each compiled check using the defined scoring procedure. Claude Opus 4.8 and Qwen3-VL 235B are also used as judges to test whether model rankings depend on the evaluator.

5.1 Main Results

GAMUT is difficult and discriminative: Gemini 3.1 Pro leads at 58.7%, and only 55.2% of its rubric elements are judged fully met. Gemini 3.1 Pro and Claude Opus 4.8 produce identical rankings, with per-model scores within 1.8 points; Qwen3-VL 235B is 4–11 points more lenient but broadly preserves the ordering. Omissions dominate failures, while weaker models also have higher contradiction rates.

The table establishes the main multimodal benchmark ranking and shows tier-level performance and verdict distributions. Gemini 3.1 Pro is highest at 58.7%, whereas Llama 3.2 11B scores 5.1%; missing checks are prevalent across models.

Performance of 14 models on 1,813 multimodal GAMUT questions, with tier scores and judge verdict distributions.
Performance of 14 models on 1,813 multimodal GAMUT questions, with tier scores and judge verdict distributions.

The three-judge comparison shows that the principal ranking is not an artifact of one evaluator. Gemini 3.1 Pro and Claude Opus 4.8 give the same ordering, while Qwen3-VL 235B assigns higher absolute scores but retains the broad ordering.

GAMUT scores for all evaluated models under Gemini 3.1 Pro, Claude Opus 4.8, and Qwen3-VL 235B judges.
GAMUT scores for all evaluated models under Gemini 3.1 Pro, Claude Opus 4.8, and Qwen3-VL 235B judges.

5.2 A Text-Only Variant

The text-only variant explicitly names each entity and removes the identification criterion that becomes trivial. Seven questions remain image-dependent and are excluded, leaving 1,806 questions; every model improves, generally by roughly 10–20 points, while the broad ordering remains.

Naming the subject in the text-only setting increases every model’s score, from +6.8 for Claude Opus 4.8 to +22.3 for GPT-4o. The broadly retained ordering supports the paper’s interpretation that visual identification is an additional cost rather than the main source of model separation.

Multimodal and text-only GAMUT scores on the same 1,806 questions, with score differences.
Multimodal and text-only GAMUT scores on the same 1,806 questions, with score differences.

6 Dataset Analysis

The dataset analysis examines whether questions are diverse and whether the structured representation is empirically necessary. It reports short but non-templated questions, domain-specific topical diversity, and rubrics whose structure, scale, and evidence support extend beyond flat fact lists.

6.1 Question Diversity

Questions have a median length of 23 words, and the 25 most common opening phrases cover under 60% of the set. Domains contain multiple recurring topics rather than a single template: plant questions span habitat, traits, care, and comparison, while vehicle questions span specifications, history, comparison, and reliability.

The stacked bars show distinct topical mixtures across the 10 domains rather than a common question template reused across categories. For example, local questions emphasize history and architecture, while food questions distribute across ingredients, history, comparison, brands, and regional variation.

Topic composition within each GAMUT domain, with row height proportional to the number of questions.
Topic composition within each GAMUT domain, with row height proportional to the number of questions.

6.2 Rubric Statistics

A question has 15.2 compiled checks on average, with median 15: 5.9 Answer-Critical, 5.8 Valuable, and 3.5 Context. Rubrics draw on about 10 web snippets per question and cite more than 9,400 distinct web pages; 97% of meta-rubric items and 94% of compiled checks have a citation.

6.3 The Structure Behind the Rubrics

Structured content is pervasive: 98% of questions require at least one component beyond Simple Knowledge, 86% use Flexible Lists, 49% Strict Lists, 17% Processes, and 6% Relationships. Flexible lists yield threshold checks and processes yield sequence checks, allowing the compiled rubric to assess open-set coverage and ordering without requiring the judge to score the original structure holistically.

The figure quantifies the empirical motivation for the two-level representation: almost every question contains nontrivial structure, most contain one or two structure types, and compiled rubrics are substantial and evidence-grounded. It reports a median of 15 binary checks and citation coverage of 97% for meta-rubrics and 94% for compiled checks.

Prevalence, co-occurrence, depth, size, and evidence grounding of GAMUT meta-rubrics and compiled rubrics.
Prevalence, co-occurrence, depth, size, and evidence grounding of GAMUT meta-rubrics and compiled rubrics.

Appendix

  • The appendix provides full prompts for question generation, selection, rubric generation and refinement, LLM judging, and text-only conversion. It also includes complete Jollof rice and Tang sancai examples, plus annotator qualification thresholds, two 45-minute rubric trainings, and 10% job auditing.

Brief Thoughts

The two-level design addresses a concrete evaluation mismatch: rich answer requirements are difficult to author as flat facts but difficult to judge holistically. Its evidence includes the widespread use of non-flat structures and stable cross-judge rankings, while the released rubrics still depend on web-sourced evidence, expert revision, and LLM-based final judgments.