Build an AI search visibility baseline you can repeat
A practical method for recording prompts, answers and citations without overstating the result.
By Michael Santiago
A screenshot can show that a brand appeared in one AI answer. It cannot show how often that happens across the questions your customers ask, or whether a later screenshot was produced under comparable conditions. A useful baseline starts by making the observation repeatable enough to inspect. You need a written method, a stable question set and a record of what actually happened.
The aim is modest: create evidence that a marketing team can revisit and discuss. You do not need a large dashboard to begin. A spreadsheet and an organized folder can support a small pilot, provided the team agrees on definitions before collecting answers. The following method is a suggested working approach, not a universal standard or a guarantee of any search outcome.
Decide which decision the baseline will support
Write a sentence explaining why you are collecting the data. Perhaps you want to learn which sources appear when people compare approaches to a problem. Perhaps management wants to know whether the company is mentioned in a defined group of product questions. These are related tasks, but they call for different summaries.
Avoid beginning with a broad score. A score forces you to combine observations before you understand them. Start with questions that someone can answer from the record: Was the company named? Was its website linked? Which other sources appeared? Did the answer describe the product accurately? Keep each observation distinct so colleagues can disagree about interpretation without losing the underlying evidence.
Build a small question set around buyer tasks
Choose one product category and identify several real tasks. A buyer might be learning terminology, comparing solution types, evaluating suppliers or checking an implementation concern. Write questions for those tasks in ordinary language. If your sales or support team has suitable examples, remove personal and confidential information before using them.
A fictional warehouse software company might test a question about choosing between inventory systems, another about migration effort and another about integrations. The questions should reflect the audience's needs rather than simply asking the engine to name your company. Include the exact wording in the record. Small wording edits can change the question being studied, so assign a version whenever you revise it.
Keep exploratory questions separate from the fixed set. Exploration is useful for discovering new themes, but adding and removing prompts during a measurement period makes comparisons harder to explain. At the end of the pilot, decide whether the fixed set needs revision and record why.
Record the collection conditions
For each observation, save the date and time, the product or interface used, the exact prompt and the language. Record relevant location or account settings that you can observe. Note whether the answer came from a consumer interface or an API, and whether it was part of an ongoing conversation. Do not quietly combine these different collection methods.
If you use an API, consult its current documentation and terms for the intended workflow. OpenAI's web search documentation describes source citation data in search-backed responses. That provides a way to retain references with an answer, but it does not make an API sample equivalent to every consumer's experience.
Store the answer in a form your team can inspect later, subject to the service's rules and your retention policy. A screenshot may preserve presentation while a text record makes comparison easier. Give each observation a stable identifier so the summary can point back to it. Never put private customer details into a public example or an unnecessary third-party prompt.
Define the events before counting them
Create separate columns for brand mention, company-site link and third-party citation about the brand. Add a short note for context: recommended, compared, criticized or merely listed. If your team needs a judgment such as “accurate description,” write down the standard and have a second person review uncertain cases.
Mark missing data carefully. A collection failure is different from a completed answer with no brand mention. An answer without links is different from a broken export that lost its links. Use explicit values such as not collected, no and yes, rather than leaving ambiguous blanks. Your future self should not have to reconstruct what an empty cell meant.
For cited pages, retain the URL and a short description of the relevant material. A citation can point to a company page, a review, a news story or a directory. Grouping these by source type may reveal a useful pattern, but keep the individual references available so the grouping can be checked.
Repeat observations with a declared schedule
Choose a schedule your team can maintain for the pilot. For example, you might collect the fixed set on several planned dates and repeat selected questions within each run. The exact cadence should match your capacity and purpose. Record it before starting, and note departures from the plan instead of filling gaps from memory.
Repeated observations help you see whether an apparent pattern is stable within your sample. They do not turn a small sample into a census of all user experiences. Report the number of completed observations alongside any percentage. “Mentioned in six of twelve collected answers” is easier to interpret than a bare visibility score of fifty.
Keep changes to the method visible. If an engine changes its interface or an integration fails, annotate the period. If you replace a prompt, compare the new version separately at first. The goal is an honest record of what you observed, including the parts that became difficult to compare.
Connect the evidence to a practical review
Suppose several answers about onboarding cite clear implementation guides from other companies. Inspect those pages and the questions that produced them. Your team may decide its own onboarding information is incomplete or difficult to find. That is a useful content discussion grounded in observed sources.
It is still a hypothesis about what to improve. The observation does not prove that publishing a similar guide will cause a future citation. Google's guidance for AI features emphasizes established SEO practices and says there are no extra technical requirements for inclusion. Keep proposed changes focused on useful, accessible information and distinguish them from promised placement.
Record each proposed action with an owner, a reason and a review date. When the action ships, add that date to the evidence log. You can then examine later observations with the change in view, while allowing for other explanations. This is more useful than attributing every movement to the most recent edit.
Publish a compact baseline report
A first report can fit into four parts: purpose, method, observations and open questions. Show the fixed question set, the collection period and the number of completed records. Include a few representative examples and link each to its evidence. Explain the limitations in the same document so the report remains understandable when forwarded.
Finish with one or two decisions the team can actually make. That might be improving a product explanation, investigating a frequently cited source or running the same set again before buying software. A tool can help once you understand the workflow you need. The baseline should first prove that your team can collect, interpret and use the evidence consistently.
