Task

"This photograph was taken inside a UK pub. Identify the exact establishment, or abstain."

Context: country: GB. Nothing else. Output: one venue claim with a disambiguating address or coordinates, or an abstention, plus confidence and the evidence used.

Tracks

  • Standard tools (primary). One harness, same text search and page reading for every model, same ceilings. No reverse-image search. We log whether search was actually used.
  • Photo only (control). Same prompt and image, no retrieval. What the model already believes about British pubs.
  • Open systems (later). Anything goes; separate table; full disclosure of every component.

Scoring

Exact-establishment accuracy against a verified venue registry. Deterministic matching first, blinded human adjudication for the rest. No fuzzy match or LLM judge is ever the sole authority.

  • "The Red Lion": ambiguous, incorrect.
  • "The Red Lion, High Street, Anytown": resolvable.
  • Chain name alone: incorrect. Nice try.

Also reported: cost per attempt and per correct answer, median and p95 latency, tool calls, abstention and error counts.

Denominators

Every attempted item counts. Abstentions, malformed output, and exhausted budgets are failures with their own status. Three runs per image; averaged within pub, then across pubs. Intervals clustered by pub, because three goes at the same Red Lion are not three Red Lions.

Data

Original candid interiors, human-verified, collected against a quota grid across UK nations, regions, and pub types. Before any model sees an image: EXIF and GPS stripped, re-encoded, hashed, deduplicated, checked for identifiable people. 50 public development pubs; 150 held out and never shown.

Trust boundary

The harness never sees the answer. Predictions are committed before a separate scorer sees the label. Model keys and ground truth live in different services.

Status

protocol
v0, not yet frozen
dataset
collecting; 0 verified pubs published
results
none
next
5 photos end to end, then a 50-pub pilot

Contribute a photo