![]() |
Author: Sven Montanus Date: 10.09.2026 Reading time: 7 min |
A growing share of B2B research no longer begins in a search engine’s results list, but in AI systems that condense several sources into a single answer. In Forrester’s study “The State of Business Buying 2026”, covering around 18,000 B2B buyers, generative AI is the most influential research source in purchasing, ahead of vendor websites and sales.1 If a brand does not appear in these answers, it is missing from part of the research before a single click occurs.
AI visibility is the measure of whether and how a brand appears in the answers of AI systems. It has two independent levels: Brand Visibility, whether the brand name is named in the answer text, and Source Visibility, whether your own website is drawn on as a source. A brand can be named without its website being cited, and vice versa. Measuring AI visibility means capturing both levels per AI system in a repeatable way.
Whether this measurement is worthwhile is hardly in question. The real question is how: which prompts, which systems, which metrics, and which measures follow from the values. As part of the GEO discipline, measurement comes first, because without a baseline no later change can be evidenced. It belongs to andweekly · Generative Engine Optimization · The Signal System™.
Acht Länder, acht unterschiedliche Marketing-Prozesse – vor dieser Situation steht Manpower. Die Folge: Uneinigkeit darüber, welche Leads Priorität haben, sowie erschwertes Benchmarking und Austausch über Best Practices.
Um internationale Vergleichbarkeit zu schaffen und Lernprozesse im Unternehmen anzuregen, will das nordeuropäische Marketing-Team um Projektleiterin Tina Hingston ein länderübergreifend konsistentes Lead Scoring und Reporting einführen. Dafür holt sie sich Unterstützung des Strategiepartners andweekly.
Von der herausfordernden und zeitaufwendigen Rekrutierung geeigneter Fachkräfte sind Unternehmen in vielen Branchen und Regionen betroffen. Das Ziel von Manpower ist es, dem Personalmangel weltweit mit innovativen Lösungen zu begegnen. Die ManpowerGroup mit Hauptsitz in den USA und Niederlassungen in rund 80 Ländern zählt zu den weltweit führenden Unternehmen in der Personalbranche.
Kerngeschäft ist die Vermittlung von Fachkräften aus zahlreichen Branchen an Unternehmen, die sich nicht mit zeitaufwendigen Rekrutierungsprozessen beschäftigen wollen. Darüber hinaus hilft Manpower, kurzfristige Personalengpässe zu überbrücken und Produktionsspitzen mit geeigneten Human Resources auf Zeit abzufedern. Zum Unternehmen gehören zahlreiche Tochterunternehmen – darunter auch der IT-Dienstleister Experis, den wir bereits bei seiner Marketing-Strategie unterstützt haben.

Die ManpowerGroup unterhält in jedem Land ein eigenes Marketing-Team, das individuelle Ansätze im Online-Marketing verfolgt. Zwar wurde HubSpot als All-in-one-Plattform für Marketing in den meisten Landesgesellschaften etabliert, doch das HubSpot-Knowhow und der hinterlegte Lead-Management-Prozess sind sehr unterschiedlich.
Das Problem bei Manpower: Die uneinheitlichen Marketing-Prozesse der Landesgesellschaften führen zu inkonsistenter Lead-Qualifizierung: Ein Lead, der in einer Landesgesellschaft als Sales Ready eingestuft wird, kann in einer anderen als Marketing Qualified Lead (MQL) eingestuft werden.
Daraus ergeben sich für Manpower folgende Herausforderungen:
Mangelnde Vergleichbarkeit. Unterschiedliche Definitionen und Prozesse machen es schwierig, die Leistung und Effektivität von Marketing-Aktivitäten zwischen verschiedenen Landesgesellschaften zu vergleichen. Ohne einheitliche Standards können sie Best Practices nicht identifizieren und erfolgreiche Strategien kaum replizieren.
Schwierigkeiten bei Zusammenarbeit und Kommunikation. Inkonsistente Definitionen führen immer wieder zu Missverständnissen und Fehlkommunikation zwischen Marketing- und Vertriebsteams, insbesondere wenn diese länderübergreifend zusammenarbeiten.
Verpasste Verkaufschancen. Unterschiedliche und nicht immer optimale Definitionen von MQLs und SQLs bewirken, dass Mitarbeitende bestimmte Leads unter- oder überschätzen. Falsche Prioritäten in der Lead-Bearbeitung kosten wiederum wertvolle Ressourcen.
Standardisierung der Marketing-Automatisierungsprozesse für eine nahtlose Customer Journey in den verschiedenen Manpower-Landesgesellschaften
Entwicklung homogener Dashboards auf globaler Ebene zur einheitlichen Erfassung, Analyse und Vergleich der Performances von Marketing-Kampagnen
Optimierung der CRM-Strategie durch Implementierung von Best Practices für Lead-Erfassung, -Qualifizierung, -Scoring und Reporting mithilfe des HubSpot Marketing Hub
Erzielung von Effizienzgewinnen durch Reduzierung von Inkonsistenzen zwischen den Landesgesellschaften
Erhöhung der Transparenz zwischen den Landesgesellschaften hinsichtlich Lead-Generierung, Lead-Qualität und Marketing-Performance zur Verbesserung der Entscheidungsfindung und Performance
The reason for this shift is a different answer format. According to G2’s report “The Answer Economy”, 51% of B2B software buyers now start their research in an AI system rather than a search engine, up from 29% a year earlier.2
In classic Google search, visibility arises through position in the results list. In an AI answer, it arises in the answer text itself. What is measured, therefore, is not rankings but two things: whether the brand is named in the text, and whether your own website is cited as a source.
These two questions are unrelated. A brand can be named without its website being cited, for instance when the answer comes from the model’s training knowledge. And a website can be cited without the brand appearing in the text. This gap between mention and citation gives rise to the brand-source gap, which later determines the measure.
How strongly an AI answer changes behaviour is shown by an analysis from Pew Research Center on Google search: when an AI overview appears, only 8% of users click a traditional search result, compared with 15% without one.3
How far mention and citation diverge also depends on the particular AI system.
AI systems do not show brand and source in the same way. Google AI Overviews usually discloses the sources it uses openly, while ChatGPT more often answers from its trained model knowledge, so a brand is named without a source being visible. An average across all systems obscures exactly these differences: the same question can produce a different result in each system.
Which systems belong in the measurement is decided by the target group. In the DACH B2B environment these are usually ChatGPT and Google AI Overviews, supplemented by Copilot or Gemini when the target group works within the Microsoft or Google ecosystem. Visibility in ChatGPT depends more on model knowledge, visibility in Google AI Overviews more on classically ranking pages.
Unlike classic SEO, there are no search-volume figures for AI systems. In SEO, search volume shows what people actually search for. For prompts in ChatGPT or Gemini no such source exists: no one knows exactly what the target group types. The prompt set is therefore not a measurement of the current state, but a reasoned assumption that presupposes a precise understanding of the target group and its buying behaviour.
For this assumption to hold, the set is built broadly and clearly structured. Two dimensions order every prompt. The first distinguishes branded and non-branded: prompts that name the brand, and generic prompts about a category, problem or solution. The second assigns each prompt to a phase of the buying journey, from early orientation through research to the decision. This shows not only whether a brand is missing, but in which phase.
The starting point is the keywords with search volume: they are the best available indication of which topics occupy the target group, and they can be translated into prompts. Personas or sales regions are added where relevant to the case. This selection is the core of any GEO work and demands market knowledge. It is also important that the set stays stable over time, otherwise the values cannot be compared across the months.
Four metrics form the core, two more provide context in terms of competition and tone. Brand Visibility shows whether the brand is named, Source Visibility whether your own website serves as a source. Within the sources, Retrievals and Citation Rate distinguish how the system treats a page: Retrievals count the pulls, Citation Rate the share of them that are actually cited in the answer.4 A page can be retrieved often and cited rarely: the system then uses it, but does not cite it as evidence.
How prominently the brand appears in an answer is shown by Position: an early mention receives more attention than one at the end. Daily values fluctuate strongly; weekly averages are meaningful. Share of Voice relates your own mentions to the competition, Sentiment shows the tone.
| Metric |
Definition | A low value indicates |
| Brand Visibility | Share of the evaluated answers in which the brand name is named | The brand is not an obvious candidate in the category for the system |
| Source Visibility | Share of answers in which a page from your own website is drawn on as a source | Content is not citable or not findable for the system |
| Retrievals | How often a page is pulled as a source, independent of citation | The page is not found as a source at all |
| Citation Rate | Share of retrievals in which the page is then cited in the answer | The content is read but not cited as evidence |
| Position | Rank at which the brand is named within an answer | The brand is named, but late and secondarily |
| Share of Voice | Share of your own brand mentions among all mentions in the competitive field | The competition occupies the topic more strongly |
| Sentiment | The tone in which the brand is described in the answer | A messaging or reputation issue, not a pure visibility issue |
What is truly revealing is the relationship between the metrics. A high Brand Visibility with a low Share of Voice points to a contested field: the brand is visible, but so are many others. A low visibility with a high Share of Voice suggests a sparsely occupied field, where visibility is easier to win. A high Brand Visibility is worth little when the Sentiment is negative. And many Retrievals with a low Citation Rate mean: the page is retrieved, but not cited as evidence.
Metrics become a report when it does three things: show development over time, separate by AI system, and mirror your own values against the competition.
For the figures to stay comparable across the months, each metric needs fixed thresholds. If visibility changes by more than 5 percentage points month over month, that is a clear swing; a Share of Voice below 15% marks a weak field, above 25% a strong one.
For data collection, andweekly uses the software Peec AI, a specialized tool for measuring AI visibility. The report leads from observation to decision and is part of andweekly · Generative Engine Optimization · The Signal System™. What remains is the last step: deriving the measure from the diagnosis.
For a metric to have an effect, a concrete optimization measure must follow from it. The path there runs through three steps: from a signal in the data, through the diagnosis, to the measure.
An example: a brand is named in ChatGPT, but its website is never cited as a source. The signal is the brand-source gap. The diagnosis: there is no page that answers the question directly and citably. The measure therefore addresses the content, not the brand.
This mapping of signal to measure is the core of GEO work. It ensures that the cause is addressed rather than the symptom, and it allows you to concentrate on the most effective measures instead of opening many fronts at once. This keeps it traceable which measure produced which change.
The article on AI content strategy shows this implementation in detail, from building citable content to sharpening the brand as an entity. Here the focus is the measurement that precedes it.
| Signal |
Diagnosis | Action |
| Brand named, website never cited |
No citable content | Content Engineering |
| Website cited, brand not named |
Brand unclear as an entity | Entity and messaging work |
| Visible in orientation, not in the decision |
No comparison or vendor pages | Content for the purchase-near phase |
| Many retrievals, few citations |
Content not robust enough | Sharpen evidence, recency and structure |
| Negative or inaccurate portrayal |
Reputation or accuracy issue | Correct the sources that shape the picture |
A visibility number is an approximation. It rests on a prompt set that models the target group’s questions, because the actual inputs in ChatGPT are not accessible to any tool. Because AI systems answer probabilistically, the same question can name different brands across two runs; individual month-to-month fluctuations are therefore often methodological, not the expression of a real market movement. Patterns become reliable only across a broad set and several weeks.
The individual metrics, too, call for careful reading. Being retrieved does not mean being cited: an analysis of 1.4 million ChatGPT prompts found that Reddit pages were retrieved en masse but cited in only around 2% of cases. In the same dataset, additional schema markup did not change citations to a statistically measurable degree.5
What matters is the position in the funnel. Visibility acts early in the buying journey, long before a lead. It is a control KPI: directly influenceable, with effect in weeks.
Traffic, engagement and conversion are impact KPIs, visible within one to three months. Leads and revenue are outcome KPIs that only show over quarters. Measuring the metric directly against leads sets a leading indicator against a lagging one that emerges only quarters later. Its contribution then appears smaller than it is.
How strongly visibility influences the pipeline cannot be captured by any tool alone. The reason lies in a typical journey: a B2B buyer first researches in ChatGPT, discovers a vendor there, then searches for it on Google and lands on the website via the organic result. In the reporting, this visit appears as organic traffic, even though the first and decisive step was an AI search. The effect of the AI search is thus attributed to another channel.
The most reliable bridge is therefore the direct question at sign-up or onboarding: “How did you hear about us?”, ideally supplemented by the query used. This self-reported attribution captures all discovery paths at once and reveals what analytics credits to the wrong channel.
AI visibility shows whether and how a company appears in the answers of AI systems: whether the brand is named (Brand Visibility) and whether your own website is cited as a source (Source Visibility). For B2B companies, this partly determines whether they appear in the AI-assisted research of their potential customers, often before a click or a contact occurs.
Being named and being cited are two separate things. In ChatGPT, a mention often stems from the model’s knowledge, while Google AI Overviews frequently discloses the sources it uses. The mechanism differs by system, which is why visibility is measured for each system separately.
Suitable tools are those built specifically for AI answers that capture mentions and citations over time. andweekly uses Peec AI for this; other specialized providers include Otterly or Rankscale. More important than the individual tool is a stable, structured prompt set and a consistent measurement methodology.
Only to a limited extent. Classic SEO tools measure rankings and keywords in the results list, that is, the position of a page in search. AI visibility, by contrast, arises in the answer text: whether a brand is named and whether a page is cited as a source. Many SEO providers have since extended their products with AI-search tracking. The underlying logic of SEO and GEO is different, however, which is why specialized tools are better suited for AI visibility. The two views complement each other but do not replace one another.
The systems use and show sources differently. Some disclose cited sources openly, others answer more from model knowledge. That is why no meaningful average can be formed across all systems: the same question can produce a different result in each system.
The sensible basis is Brand Visibility and Source Visibility, supplemented by Retrievals, Citation Rate and Position for the source and prominence level, and by Share of Voice and Sentiment for competition and tone. Taken individually, each metric misleads; together they produce a reliable picture of your own visibility in AI systems.
Measuring AI visibility replaces guesswork with a foundation: a structured prompt set along the buying journey, clear metrics and separate consideration per AI system. It shows in which phase and in which system a brand is missing, and thus provides the starting point for all further work on visibility. At andweekly, part of Generative Engine Optimization, it belongs to andweekly · Generative Engine Optimization · The Signal System™.
Further analysis of visibility in AI systems appears in the Perspektiven newsletter: Subscribe to the Perspektiven newsletter.
1 Forrester Research (2026): The State of Business Buying
2 G2 (2026): The Answer Economy – AI Search Insight Report
3 Pew Research Center (2025): Google users are less likely to click on links when an AI summary appears in the results
4 Peec AI (2026): Metrics overview (Brand Visibility, Source Visibility, Retrievals, Citation Rate)
5 Search Engine Journal (2026): Brands Are Tracking AI Visibility, But Are They Measuring The Right Things?