Key takeaways
- Bing’s AI Performance report in Bing Webmaster Tools shows how often your site is cited in Bing Chat, Copilot and partner tools, even if you get almost no normal search traffic.
- Grounding queries are the real searches LLMs fire behind the scenes to support a prompt, and they often use language that humans never type into Google.
- You can spot repeated patterns in these grounding queries, then create focused pages and H2 sections around them to earn far more AI citations across many LLMs.
- This opens both good and bad tactics: you can shape how tools describe your brand, your competitors, and even whole markets, with quite small sites.
If you only care about the bottom line, here it is: you can use Bing’s AI Performance data to see the weird queries LLMs use behind the scenes, then build short, focused pages around those terms to get a disproportionate share of citations in AI answers, even if your site is tiny.
That means more visibility in Bing Chat, Copilot, and any tool that leans on them, plus indirect gains in Google and other LLMs that copy similar patterns, without writing 3,000 word guides for every topic.

What Bing’s AI Performance tab really shows you
Bing’s AI Performance report looks simple at first glance, but it is quietly exposing a side of search we never had access to before.
And I think a lot of marketers are looking at it, shrugging, and going back to traditional impressions and clicks, which is a mistake.
Clicks vs citations: two very different worlds
Let me give you a pattern that keeps popping up: tiny site, hardly any traffic, and thousands of AI citations.
I tested this on a side project in a boring B2B niche where search volumes are low and no one brags about traffic screenshots on social media.
| Metric (last 90 days) | Traditional search | AI Performance |
|---|---|---|
| Organic clicks (Bing) | 80 | Not shown |
| Organic impressions (Bing) | 3,200 | Not shown |
| AI citations | Not in normal report | 34,000+ |
| Pages with most citations | Generic blog posts | Short, very focused comparison pages |
So the site was basically invisible in normal Bing search, but it was getting cited thousands of times in AI answers inside Copilot and partner tools.
Once you see that once, you stop treating AI traffic as some small side channel and start treating it as a parallel index with its own logic.
Most people are trying to grow search traffic, while ignoring that LLMs are quietly quoting sites that get almost no clicks.
Where these citations actually come from
Bing groups AI activity from a few places in that report.
The wording changes slightly over time, but the buckets tend to include:
- Bing Chat and Copilot inside Bing search
- AI answers packed into normal Bing results
- Copilot powered experiences in Edge and Windows
- External tools that are whitelabeling Bing or using its API as a backend
This last group is what I think most people underestimate.
A lot of SaaS dashboards, “AI research tools”, and brand monitoring products are leaning on Bing or similar APIs to fetch web data, then repackage it as something fancier.
From the outside you see a polished report about your brand.
From the inside, it is often just a handful of grounding queries silently hitting the web, grabbing a few pages, and then letting an LLM talk about them.
Why AI Performance is different from Search Performance
Traditional search data answers a simple question: which queries did humans type and where did your pages show.
AI Performance is about something else: which prompts triggered LLMs to look for data, which web queries that produced, and when your pages got pulled into that process.
That means the metrics look and behave differently:
| Feature | Search Performance | AI Performance |
|---|---|---|
| Main user action measured | Human typing a query | Prompt sent to an LLM |
| Key metric | Clicks, impressions, position | Citations (and related queries) |
| Surface | Classic search results page | AI answers, summaries, chat |
| Language style | Human phrasing | Machine generated subqueries |
So if you open AI Performance and try to treat it like regular keyword data, it will feel strange.
You are not wrong to feel that way; the source is different.

Grounding queries: what they are and why SEO people should care
Grounding queries sound like something very academic, but the idea is simple.
When someone types a prompt into an LLM, the model often turns that prompt into one or more web searches to “ground” its answer in fresh data.
The prompt is not the query
This is what confuses many SEOs: they look at a long prompt and assume it is the same as a search keyword.
Most of the time it is not.
Imagine a user types this into Copilot:
“Compare three vendors for enterprise payroll compliance in Europe and explain which one has stronger monitoring and reporting features.”
That single prompt might trigger internal searches like:
- “enterprise payroll compliance platforms europe comparison”
- “payroll compliance vendor monitoring reporting evaluation”
- “global payroll compliance tool feature checklist”
Those are the grounding queries.
They are the bridge between the natural language prompt and the pages the model will pull in and cite.
The prompt belongs to the user, the grounding query belongs to the LLM.
Once you think about it that way, a lot of odd looking keywords in your reports start to make sense.
Why grounding queries look so strange
When I first looked at the AI Performance report on that small B2B site, one thing jumped out immediately.
I kept seeing words that normal users rarely type, floating alone or glued to generic phrases.
Things like:
- “evaluate”
- “assessment”
- “scoring criteria”
- “feature comparison matrix”
I do not usually see a human go to Bing and type just “evaluate”.
But I can imagine a SaaS platform asking an LLM to “evaluate provider X” and then leaving the rest to the system.
So the LLM takes “evaluate” and the brand or topic, bolts on some generic modifiers, and turns that into grounding queries.
That is why your reports suddenly show odd phrases that rank very high, get zero clicks, and still generate thousands of AI citations.
Why this matters for rankings beyond Bing
Here is the part I think some people are still underestimating.
LLMs across different products tend to reuse similar search patterns, because they are solving the same problems.
Perplexity, ChatGPT with browsing, Gemini, Copilot, even smaller niche tools, all have to figure out: “What should I search for on the web to answer this prompt safely.”
That usually leads to a small set of recurring terms around evaluation, monitoring, scoring, reviews, pros and cons, and so on.
If you rank for the common wording in one LLM’s grounding queries, there is a good chance you will fit cleanly into others too.
I am not saying the queries are identical, they are not.
But the overlap is strong enough that when I tuned pages for one set of grounding phrases, I started seeing the same pages surface in other AI tools that have nothing to do with Bing officially.

A simple workflow to mine and use grounding queries
So how do you turn this from theory into actual content changes that drive more AI visibility.
Let me walk through a workflow that I have been testing, which is not perfect, but it is practical.
Step 1: Pull and sort your AI Performance data
First, go into Bing Webmaster Tools, open the AI Performance tab, and export the data for the last 90 days.
I like 90 days because it is long enough to see patterns, but not so long that you mix very different model behaviors.
Then I sort by citations, descending.
You will usually notice that a few pages and a few grounding queries dominate.
- Find the top 10 grounding queries by citations.
- Group them loosely by theme or repeated wording.
- Note which pages they map to.
At this stage, do not overthink it.
You just want to see which odd phrases are doing a lot of heavy lifting.
Step 2: Look for repeated “AI-ish” language
When you scan those top queries, look for words you do not normally target in classic keyword research.
In my experiments and client accounts, things like these show up a lot:
- evaluate / evaluation
- assess / assessment
- benchmark
- scorecard
- capabilities overview
- feature monitoring
- sentiment summary
These are words that sound like they came out of a product manager meeting, not a real searchers mind.
Which is exactly what you would expect if tools are driving a big chunk of those prompts.
Grounding queries often expose the lazy prompts that SaaS tools send to LLMs behind the scenes.
Once you spot those repeated patterns, write them down in a simple table.
| Grounding term | Example full query | Page currently ranking | Citations (90 days) |
|---|---|---|---|
| evaluation | “global payroll platform evaluation” | /global-payroll-platforms | 12,400 |
| benchmark | “hr compliance tools benchmark” | /hr-compliance-tools-table | 6,900 |
| scorecard | “vendor risk management scorecard” | /vendor-risk-basics | 3,100 |
Now you have a short list of “AI words” that LLMs already like to pin to your niche.
These become your starting points.
Step 3: Decide when to edit a page vs create a new one
This is where I disagree with some SEOs who want to stuff everything into existing pages.
If a grounding query is close to a page topic you already have, adjust that page, but if it is quite different, create a new asset.
For example:
- If your page is “Payroll compliance in France” and you see “France payroll compliance evaluation” cited, you can add an H2 like “Payroll compliance evaluation for France” and a short section explaining how to judge providers.
- If your grounding query is “benchmark payroll compliance tools” and your current page is just a definition of payroll compliance, that is too far; create a separate comparison page.
My rough rule is this: if the intent feels like “compare or judge providers”, that deserves its own page most of the time.
A short, focused one works well; you do not need 2,000 words of history and philosophy.
Step 4: Structure content in a way LLMs can quote cleanly
Once you know which grounding phrases matter, you can build pages or sections that are easy for LLMs to lift from.
I like a simple layout:
- H2 with the grounding phrase used naturally near the start
- A short intro sentence that restates the topic in plain language
- A table that compares options, features, or criteria
- A short paragraph that interprets the table
For example, for “global payroll platform evaluation”, a page might start like this:
Sample structure for an evaluation page
Below is a simple outline that has worked well for me in similar cases.
H2: Global payroll platform evaluation
A one sentence summary of what a buyer should care about.
H3: Key evaluation criteria
- Coverage by country
- Compliance updates
- Reporting depth
- Support model
H3: Comparison table
| Platform | Countries covered | Compliance updates | Reporting | Support |
|---|---|---|---|---|
| Vendor A | 120+ | Automatic, weekly | Custom dashboards | 24/7 chat and phone |
| Vendor B | 60+ | Manual, monthly | Basic exports | Email only |
| Your product | 90+ | Automatic, daily | Drillable reports | Dedicated manager |
Then a short paragraph that clearly states who each platform is best for.
LLMs love this sort of clean structure, because they can quote a line, quote a row, or summarize the whole thing.

Using grounding queries to shape how LLMs talk about brands
Now we get into the part that feels slightly uncomfortable, but also very real: you can use these queries to influence how tools describe brands, including your own and competitors.
I am not talking about lying or making up facts, but I am also not going to pretend this cannot be abused.
Brand evaluation prompts are low effort on the tool side
Many “AI brand monitoring” dashboards work roughly like this:
- User selects a brand or enters a URL.
- Tool sends a prompt like “Evaluate the perception and online presence of Brand X.”
- LLM fires grounding queries with words like “evaluate”, “rating”, “reviews”, “sentiment”.
- Tool wraps the answer into charts and a PDF.
Notice what is missing: clear criteria, defined data sources, narrow scope.
Because the prompt is so vague, the LLM has a lot of freedom to pick whichever pages it finds first and trusts most for that pattern of words.
When the prompts are lazy, the web pages that match those vague grounding queries gain a lot of power.
If that sounds fragile, it is.
But from an SEO point of view, fragility can be an opportunity if you handle it with some ethics.
Ways brands can responsibly influence AI evaluations
Let me run through a few ways a normal company could use this without crossing lines.
You can judge where your own red lines sit, of course.
1. Create honest evaluation pages that include your brand
Suppose there is a grounding query like “cloud backup providers evaluation” that often pulls in pages you do not control.
You can create your own page with that phrase in the H2, and include:
- A clear explanation of how to evaluate cloud backup providers.
- A simple comparison table with key features.
- A neutral summary of strengths and tradeoffs for each vendor.
You do not have to attack anyone; just make the evaluation more complete than what currently ranks.
In my tests, LLMs start to grab this sort of page very fast, often within days, because it fits the grounding query perfectly.
2. Add “evaluation” sections to existing high trust pages
If you already rank well for “what is X” and “how X works”, you can bolt an evaluation section onto those pages using the new wording.
For example:
- Add an H2 like “How to evaluate X platforms” near the middle of the page.
- List 4 or 5 criteria buyers should use.
- Explain in 2 or 3 sentences how your product performs on those criteria.
This often lets you catch both regular search queries and grounding queries, without needing a separate page.
I still lean toward separate pages when the topic is quite commercial, but I know some brands prefer fewer URLs.
3. Fix unfair or outdated narratives via reputation pages
Sometimes AI answers keep repeating the same old criticism or outdated comparison, because the only pages that mention it are years old.
In that case, a focused page with wording close to the grounding queries can help.
Think of something like “Brand X pricing review” or “Brand X outage history evaluation” that shows up in your AI Performance data.
You can write a page that:
- Honestly explains what changed since older reviews.
- Shows recent data, charts, or public status reports.
- Clarifies what is true and what was a one off incident.
LLMs are not great at time awareness, but they do like fresh, well structured content that speaks directly to the query.
So this sort of page often gets cited when the model tries to talk about your brand history.
A gray area: influencing how competitors are framed
Now we come to the part where I think a lot of SEOs will quietly experiment, even if they do not say it publicly.
Grounding queries often ask about “top providers”, “best tools”, or “vendor comparisons” in a niche.
If you rank for those queries with your own comparison pages, you gain leverage over which competitors are mentioned and how.
Here is where you have to be careful with claims, but there is room to frame things:
- You can choose which competitors to include.
- You can focus the criteria on features where you are stronger.
- You can state that some tools do not support certain use cases, as long as that is true.
You should not lie about competitors, but you also do not have to present their marketing in the best possible light.
In practice, when I did this for a developer tooling client, we saw their product start to appear more often in high intent AI answers inside code assistants that pull from the web.
The grounding queries were not public, but the pattern was clear from logs and user feedback.
How fragile is this tactic over time
I do not think this sort of “LLM SEO” is stable forever.
Models change, prompts inside tools change, and the wording of grounding queries will shift.
So if you are thinking of this as some magic evergreen trick, you are overestimating it.
I see it more like this:
- Short term edge: grounding queries are like new keywords without competition.
- Medium term: competitors see the same data and start to target similar language.
- Long term: vendors tighten their prompts, and the easy manipulation fades.
But during that early phase, the lift can be huge compared to the work involved.
For one site I worked on, a single 500 word evaluation page drove more AI citations in three months than a 40 post blog had done in a year.
Practical tips to avoid overreacting to grounding queries
I should also say this, because some people are now rewriting entire sites around one or two words from these reports, which is overkill.
Grounding queries are useful, but not sacred.
Do not throw away normal keyword research
You still need pages that match the phrases humans type.
“Best HR software for startups” still matters, even if the LLM grounding query says “hr software vendors evaluation”.
So I tend to:
- Start from classic keyword research for core pages.
- Use grounding queries to create extra layers: evaluation, benchmarks, scorecards, comparisons.
- Leave highly stable evergreen content mostly alone.
When people tell me they want to rewrite high traffic pages just to inject the word “evaluation” three times, I usually push back.
You can add new pages and internal links instead of risking cannibalization.
Track AI citations like a separate funnel
Another mistake I see is people staring at citation numbers without asking what those citations do for the business.
That can lead you to chase whatever queries produce flashy counts, even if they attract the wrong audience.
What I have been doing with clients is simple:
- Create a basic dashboard that tracks AI citations by page.
- Overlay that with assisted conversions, demo requests, or signups by landing page.
- Watch which AI cited pages feed useful visits later through branded searches or direct traffic.
The correlation is not perfect, but you start to see which types of evaluation content actually move people closer to buying.
That helps you avoid chasing ego metrics like “100k citations” that do not lead to revenue.
Accept that some of this will break
I am always a little nervous when a tactic depends heavily on a single vendors interface choice.
Bing may change how much AI data they expose, or how they count citations, or which partners are included.
So I would treat this as a boost, not a foundation.
Your foundation is still:
- Pages that answer clear problems better than alternatives.
- Technical health that keeps crawling and indexing clean.
- Links and mentions that prove people care about your content.
Grounding queries sit on top of that, like a layer that lets you intercept a new stream of machine triggered “searches”.
Useful, but not your entire strategy.

How to get started this week without overcomplicating it
If you want a simple starting plan, something you could follow over the next week or two, here is how I would approach it.
You do not need a big team or a giant domain to try this.
Day 1-2: Pull data and spot patterns
- Open Bing Webmaster Tools and export 90 days of AI Performance data.
- Sort by citations and highlight the top 20 grounding queries.
- Circle any odd words that keep repeating: “evaluation”, “benchmark”, “scorecard”, “monitoring”, and so on.
You should end this step with a short list of 3 to 5 “AI-style” phrases that matter for your site.
If you have nothing at all yet, it might point to a different problem: you are not ranking enough in Bing for LLMs to notice you, which is a separate issue to fix.
Day 3-5: Create or adjust 2-3 pages
- Pick one theme, like “platform evaluation” or “vendor benchmark”.
- Create a new page focused on that theme, with a table and clear criteria.
- Optionally, add a short H2 section about evaluation to one existing page that already ranks.
Do not worry about making the “perfect” page.
It is better to publish a clear, honest, 400-600 word comparison page than to sit on a draft for a month.
Day 6-7: Internal links and measurement
- Link to your new pages from related content using simple anchor text like “our X platform evaluation”.
- Set up a small report where you track citations and any conversions tied to those pages.
- Mark the publish date so you can compare AI Performance data in a month.
Then give it some time.
Bing has to crawl the pages, LLMs have to start bouncing into them during grounding, and tools have to run their prompts.
Treat this as an ongoing test: a way to learn how LLMs see your niche, not a one-off trick that magically fixes growth.
Some pages will do nothing, some will get a few citations, and every so often one short page will suddenly become the most quoted thing you ever wrote.
That unevenness is normal; it is how experimentation usually looks under the hood.
Where this fits in your broader SEO approach
The honest truth is that AI grounding queries do not replace traditional SEO, but they do change how far a small site can punch above its weight in certain areas.
A 500 word evaluation page on a low authority domain is not going to outrank a giant site for “best software” in classic Google, but it can quietly become the default citation in a niche AI tool that thousands of decision makers use.
That gap between classic rankings and AI citations is where the leverage lives.
If you keep showing up in the boring, machine generated queries that other marketers ignore, you give LLMs fewer excuses to hallucinate random competitors into the story.
You will not control everything, and you will be wrong in parts of your approach, same as everyone else right now.
But that is fine; the teams who experiment with this data, question their own assumptions, and keep writing clear, honest evaluation content are the ones who will quietly shape how AI tools talk about their markets for the next few years.
Need a quick summary of this article? Choose your favorite AI tool below:


