A model predicts text
An LLM is a probabilistic text generator. Small differences in prompts, settings, and model versions can change outputs.
Loading…
Learn
Large language models are probabilistic text generators: ChatGPT and Gemini may retrieve public web signals, summarize them, and change answers when prompts, models, or dates differ — which is why Coastline Metrics stores prompt, model, and date on every baseline and verification rerun.
Large language models are probabilistic text generators: ChatGPT and Gemini may retrieve public web signals, summarize them, and change answers when prompts, models, or dates differ — which is why Coastline Metrics stores prompt, model, and date on every baseline and verification rerun.
An LLM is a probabilistic text generator. Small differences in prompts, settings, and model versions can change outputs.
Many products add retrieval so the model can quote current web sources. Missing or inconsistent site data may not be retrieved or cited.
Even when a UI shows citations, the model may summarize or blend sources. You still need to verify what was used and what changed.
To compare runs, hold constant: prompt pack, locale, provider, model, and date. Otherwise differences may be drift.
A buyer may ask an assistant for a provider that offers a specific service in a specific market. The answer can combine model knowledge with retrieved webpages, profiles, directories, and other public sources. If those sources disagree about the service, location, or business identity, the assistant may omit the business or describe it incorrectly.
That is why useful AI visibility work starts with public facts the business can verify—not assumptions about what a model should know.
The same question can produce a different answer when wording, product mode, retrieval tools, model version, or date changes. A single screenshot documents one observation; it does not establish a stable trend.
A defensible comparison preserves the question set and relevant run conditions, stores the returned evidence, and labels any mismatch before describing movement.