The unit you compare should be an experience and a question, not just a model label.
Name the surface you tested
A model family, consumer assistant, API configuration, search mode, and coding agent are not interchangeable test surfaces. Record exactly what your collection uses and which capabilities are enabled. If you cannot identify the underlying model version, state that limitation rather than assigning a plausible name.
Keep market, language, account context, and conversation state in the test record where available. A comparison becomes difficult to interpret if one sample uses a new conversation while another carries several earlier preferences.
Use a shared question panel
Run the same baseline questions across the experiences your audience uses. Preserve brand names and decision constraints during translation, with native-language review for important markets. Document unavailable features and failed collections separately.
For each surface, record whether an answer was produced, whether it included sources, and whether it used an ordered recommendation. Some measures apply only to some answer types. Do not manufacture a rank for an answer that never presented a list.
Read disagreements as research
When two experiences recommend different vendors, inspect the descriptions and sources. One may emphasize cost; another may prioritize integrations. The disagreement can reveal an information gap or an audience assumption worth testing. It does not automatically establish which answer is correct.
Compare at topic level before aggregating. A brand can be strong on implementation questions and weak on category discovery. A single cross-platform average hides that difference and may direct the team toward the wrong content.
Report within the sample’s limits
A controlled prompt panel describes performance on that panel. It does not establish platform market share, real user demand, or all possible personalized answers. Those questions require other evidence. Keep measurement findings and audience assumptions visibly separate.
Give the reader a comparison table with question group, surface, observations, brand visibility, citation presence, and collection notes. Include example answers. The table should make the conditions understandable before anyone interprets the winning column.





