1 · Question bankA curated buyer-intent benchmark
Each industry gets a curated bank of standardized buyer questions covering discovery, DIY-vs-hire, vetting, pricing, comparison, local availability, urgency and red-flag scenarios. Banks are frozen within the edition; the exact questions are published on each industry page.
2 · CollectionOne response per model per question
The collection contract expects one API response per question and model (chatgpt: gpt-5-mini · claude: claude-sonnet-5 · gemini: gemini-3-flash-preview), using identical question wording within the edition window. Published pages disclose observed response counts against that expectation rather than silently treating missing responses as complete.
3 · CodingA fixed 12-behavior rubric
Responses are classified against the same 12 binary behaviors (definitions below). Per-model percentages are simple proportions over observed responses. The displayed equal-model mean gives each available model the same weight.
4 · QAAutomated integrity checks before publication
The import and test pipeline validates required fields, response coverage, percentage bounds and aggregate consistency. Missing responses are reported. This repository does not presently document double coding, reviewer counts or adjudication, so no human-review claim is made.
5 · EditionsDated snapshots on stable URLs
Each file names its edition and collection date. A future edition can update the stable page after passing the same checks, but no refresh cadence or historical trend is claimed until multiple retained editions exist.
Audit trailPublicly inspectable measurement fields
Published study files expose the question bank, API model IDs, observed response counts, per-model behavior rates, question-level agreement and question-level disagreement. Raw answer text is not part of the current public dataset claim.