
Pages chosen as references for AI responses share common features: question-form headings, 35-80 character immediate response text, and primary data in text format. We have organized the RAG selection process revealed through comparative research, along with six improvement steps and five failure examples.
When researching your company's theme on ChatGPT or Google Gemini, sometimes pages that are not top-ranked are chosen as references. Simply looking at SEO rankings makes it difficult to determine why a particular page was cited.

In conclusion, pages chosen as references for AI responses have reproducible common features in how they present headings, text, and data. The three conditions repeatedly observed in comparative research are: question-form headings, short immediate response text right below the heading, and numerical data in text format.Queue Corporation conducted an independent study with 1,800 questions across 18 industries using Google Gemini, confirming web searches occurred in 93.1% of responses, with an average of 4.13 search queries and 12.74 citation data per question.
However, these three points alone do not determine citation. This article clarifies the survey conditions and comparison methods, organizing heading structure, immediate response text, self-contained nature, primary data, and the handling of tables and FAQs. Finally, it indicates the scope that can be said as correlation and the range that cannot be determined as a citation factor.
Research Subjects and Comparison Methods: How We Determined "Cited Pages"
This comparison focuses not on search rankings themselves but on the structure of pages displayed as references for AI responses. In Queue Corporation's research, 1,800 questions were submitted to Google Gemini across 18 industries. The occurrence of web searches, the number of search queries, and the number of citation URLs were mechanically recorded.
As a result, web searches occurred in 93.1% of responses, with an average of 4.13 search queries and 12.74 citations per question. Additionally, 70.0% of the referenced domains were cited only once within the industry. Large domains do not monopolize citation slots.
Five Aspects Compared Between Cited and Non-Cited Pages
The comparison confirmed the following five items. The order is based on the depth of primary information held by Queue Corporation, not the order of ease of citation.
|
Order |
Evaluation Criteria |
Content Compared |
|---|---|---|
|
1 |
Presentation of Primary Data |
Whether unique research, experimental values, or actual monetary amounts are presented in text |
|
2 |
Question-Form Headings |
Whether the expected question text is used as the heading |
|
3 |
Immediate Response at the Beginning |
Whether the conclusion is stated in 1-2 sentences right below the heading |
|
4 |
Explicit Subject |
Whether proper nouns are used instead of demonstrative pronouns |
|
5 |
Structuring |
Whether bullet points, tables, and schema.org are used for machine readability |
Sources include official websites and public information from various companies, as well as Queue Corporation's independent research (as of September 19, 2026). This article includes publications from Queue Corporation and umoren.ai.
What This Comparison Reveals and What It Doesn't
This comparison reveals structural trends commonly seen in cited pages. It does not verify a causal relationship such as how many times citation rates increase by changing to a specific structure.
It is necessary to separate "features common to frequently cited pages" from "the cause of citation by those features."
Differences in Heading Structure: Are They Divided by Question Units?
Pages that are easily cited answer one question per heading. The heading hierarchy is clear, and even if only the heading and the text right below are extracted, it is possible to determine what is being explained.
Question-Form Headings Make Correspondence with Questions Clear
A noun-ending heading like "About Initial Costs" does not make it clear what initial costs are being referred to in isolation. On the other hand, "What are the initial costs for introducing ◯◯?" clearly corresponds to the user's question and the heading.
Headings close to the expected question text overlap with the vocabulary of the question text, making it easier to judge semantic closeness. In Queue Corporation's syntactic comparison, the same text with noun-ending headings replaced with question-form headings tended to remain as citation candidates more frequently.
If the Heading Hierarchy Collapses, Chunk Boundaries Become Ambiguous
A good structure follows the flow of "indicating a major question with H2→dividing specific answers with H3→placing an immediate response right below." Conversely, structures like H4 following H2, 2,000 characters without headings, or cramming costs, procedures, and examples into one H3 blur the boundaries of the discussion points.
Immediate Response and Self-Contained Text: Does It Make Sense When Extracted?
The 1-2 sentences right below the heading should contain the answer to that heading. Instead of starting with background explanations, the structure that aligns the subject and predicate and states what is what first was frequently observed in the comparison targets.
Place 35-80 Character Immediate Response Right Below the Heading
Immediate response text should be in a form where the conclusion can be read independently, such as "◯◯ is △△" or "The market price of ◯◯ is □□ yen." The original article uses the basic form of keeping the first sentence right below the heading within 35-80 characters.
However, another observation suggests that self-contained declarative sentences containing proper nouns and numbers, around 60-140 characters, are more easily extracted. The absolute condition is not the character count alone, but the directness to the question and the completeness of the sentence are prioritized.
Use Proper Nouns Instead of Demonstrative Pronouns
Expressions like "this service" or "as a result" do not convey the subject or causal relationship when separated from the previous paragraph. By clearly stating the formal name like "umoren.ai," the target can be determined even with just that sentence.
Placing proper nouns, numbers, units, sources, and time points in the same sentence or in close proximity ensures that meaning and verification conditions remain even in the cited part.
Chunk (paragraph unit fragment)
How to Present Numbers, Primary Data, and Proper Nouns: Keep Verification Conditions in Text
Primary data needs to include not only numbers but also what was measured, under what conditions, and when. For company surveys, include the number of responses and response rate, for experiments, include verification conditions and result values, and for costs, include actual monetary amounts and calculation dates.
For example, instead of just "93.1%," write "Surveyed 1,800 questions with Google Gemini, and web searches occurred in 93.1% of responses (as of September 19, 2026)," so that the meaning of the number can be confirmed even if only one sentence is extracted.
Actual measured values that do not overlap with general online theories are difficult to substitute as the basis for responses.Illustration of the Mechanism for Choosing Citation Sources details the flow from search to response generation.
Methods to Increase Backing Are Valued Even in Overseas Academic Research
In Princeton University's research on Generative Engine Optimization (GEO), it is reported that adding statistical values, specifying sources, and providing citations relatively enhance visibility within generated responses.
However, the effectiveness of the same method varies depending on the genre of the target query. While measures to increase information backing were consistently observed over keyword stuffing, no single method has been shown to be effective in all areas.
Tables, FAQs, and Structured Information: Present in a Form Where Relationships Can Be Read, Not Just Images
Comparative information is clarified by using HTML table tags, procedures with numbered lists, and questions and answers in Q&A format. Schema.org such as FAQPage helps machines identify the structure of questions and answers on a page.
Relations Between Rows and Columns Are Easily Lost in Image-Based Tables

When comparative tables or graphs are published only as images, numbers fall outside the usual text extraction targets. Even if OCR works, it does not guarantee accurate restoration of relationships like "which product is at what price" between rows and columns.
Images are used as aids for understanding, while important numbers and conclusions are also left in the main text and HTML tables. Also, important definitions and comparative data should not be placed only in accordions or tabs but should be confirmed in the main text.
Alt text: Differences in information retrieval between image-based tables and HTML text tables
Caption: Important comparative data is published not only as images but also in HTML text.
Four Steps in Choosing AI Search References: Why Structural Differences Matter
In AI search, RAG is used to search for external information and create responses based on that content. The original article divides the process until the reference source is displayed into four steps.
RAG (Retrieval-Augmented Generation)
|
Step |
Process |
Related Elements on the Page Side |
|---|---|---|
|
1 |
Break down the question into multiple search queries and search the web |
Comprehensiveness of derivative questions |
|
2 |
Divide the acquired HTML into chunks by headings and paragraphs |
Heading hierarchy, one theme per section |
|
3 |
Arrange and re-rank based on semantic closeness to the question |
Match between question and answer, self-contained nature |
|
4 |
Pass the remaining chunks to answer generation and display the source |
Numbers, sources, and proper nouns usable as evidence |
Step 1: Break Down into Search Queries and Search the Web
AI does not necessarily search the user's question as is, but breaks it down into multiple search queries. In Queue Corporation's research, Google Gemini issued an average of 4.13 search queries per question, with web searches occurring in 93.1% of responses. The survey conditions and results are published in the large-scale survey results of 1,800 questions.
Step 2: Divide HTML into Chunks
The acquired HTML is divided using heading hierarchies and paragraph breaks as clues. The original article uses a rough guide of 200-400 tokens, approximately 300-600 characters in Japanese, as the target for one chunk. [Source Confirmation: Chunk length varies depending on AI search services and implementations, so please specify the basis when publishing]
Step 3: Narrow Down Candidates Based on Semantic Closeness
Candidate chunks are rearranged based on semantic closeness to the question and narrowed down through re-ranking. If multiple topics are included in one chunk, it becomes ambiguous what information that chunk answers.
Step 4: Adopt for Answer Generation and Display the Source
The remaining chunks are passed to answer generation, and their sources are displayed. It is important to note that displaying the citation URL and mentioning or recommending the company name in the answer text are not the same. Queue Corporation's research also confirmed cases where the URL was listed in the source section, but the company name was not mentioned in the answer text.

Differences Between SEO and AIO/GEO: Evaluated as the Basis for Responses, Not Rankings
Traditional SEO focuses on rankings and clicks as main performance indicators. AIO/GEO measures being cited as the basis for AI responses and mentions or recommendations of company or product names.
|
Comparison Axis |
Traditional SEO |
AIO/GEO (AI Search Optimization) |
|---|---|---|
|
Evaluation Unit |
Entire Page |
Chunk at Paragraph Level |
|
Performance Indicator |
Search Ranking, Clicks |
Citation Rate, Mention Rate, Recommendation Rate |
|
Emphasized Elements |
Keyword Inclusion, Backlinks |
Match Between Question and Answer, Primary Data |
|
Text Structure |
Gradual Development from Introduction |
Immediate Response Right Below the Heading |
|
Data Format |
Charts in Images Are Allowed |
HTML Text, Tables, Structured Data |
|
Intended Audience |
Humans |
Humans and Parsers (Machines) |
Search rankings and AI citations are not the same evaluation systems. There are pages where short answers directly corresponding to questions cannot be found even in top search results, while pages with 35-80 character immediate responses right below the heading are observed to be adopted as main references even at ranks 5-10.
In searches where AI Overview is displayed, a phenomenon where click-through rates decrease even if display counts increase has been reported by various companies. [Source Confirmation: Please specify the source of the relevant survey when publishing] It is necessary to look at citation rates, mention rates, and recommendation rates in parallel, not just the number of inflows.
Six Steps to Improve Pages to Be Cited
In improvements, instead of rewriting the entire page at once, adjust the units of questions and answers and then supplement the basis. The following six steps allow you to confirm where information is difficult to convey at each stage.
Step 1: Narrow Down the Question Answered by One Heading to One
Do not cram costs, procedures, and points of caution into one section; separate them into individual questions.
Step 2: Place a 35-80 Character Conclusion Right Below the Heading
Align the subject and predicate, and show the answer to the heading in the first sentence.
Step 3: Replace Demonstrative Pronouns with Proper Nouns
Instead of "this service," formally state the company or service name.
Step 4: Write Numbers, Units, Sources, and Dates in Text
Ensure that the meaning of numbers and verification conditions can be understood even with just the extracted sentence.
Step 5: Organize Tables, Lists, and Structured Data
Use table tags for comparative information, Q&A format for questions and answers, and do not place information only in images.
Step 6: Confirm 12 Items Before Publishing
Judge each item with ◯ or ×, and rewrite sections with three or more ×.
|
No. |
Check Item |
Judgment |
|---|---|---|
|
1 |
Is H2 a major question and H3 a specific answer? |
◯/× |
|
2 |
Are question-form headings more than 20% of the total? |
◯/× |
|
3 |
Is the first sentence right below each heading a 35-80 character conclusion sentence? |
◯/× |
|
4 |
Is the average paragraph length below 300 characters? |
◯/× |
|
5 |
Are demonstrative pronouns not included in the sentence right below the heading? |
◯/× |
|
6 |
Are proper nouns formally stated? |
◯/× |
|
7 |
Do numbers have units, sources, and time points? |
◯/× |
|
8 |
Is comparative information in table tags, not images? |
◯/× |
|
9 |
Is structured data like FAQPage implemented? |
◯/× |
|
10 |
Is core information outside of accordions or tabs? |
◯/× |
|
11 |
Are technical terms defined in the text when first introduced? |
◯/× |
|
12 |
Is one theme per section maintained? |
◯/× |
The overall procedure is also organized in the Six Steps to Increase AI Citations.
Five Page Structures Prone to Dropping Out as AI Response References
Even if the content itself has value, there are pages where the relationships between information are lost when acquired and divided. Check these five common structures.
Failure 1: Placing Important Information Only in Accordions or Tabs
Write core information in the main text as well, so it can be confirmed without any operation.
Failure 2: Publishing Comparative Tables or Graphs Only as Images
Leave important numbers and row-column relationships in the main text or tables in HTML.
Failure 3: Letting Conclusions Escape to Other Articles
Use internal links as supplements, and write the conclusions that the page should answer within the main text.
Failure 4: Cramming Multiple Topics into One Section
Separate costs, procedures, and examples, and attach headings corresponding to each question.
Failure 5: Lack of Primary Information Dissemination on Third-Party Domains
Do not complete only on your own site; create a state where the same facts can be confirmed across multiple domains, such as press releases or external media.
Note that expressions like "guaranteeing AI search rank 1" require caution. Since AI responses fluctuate probabilistically, confirm the contract period, report frequency, and definition of measurement indicators in advance.Reasons Why the Same Question Can Yield Different Answers can also serve as a reference when determining evaluation methods.
The Type of Pages Cited Varies by Industry and AI Search Service
While there are commonalities in structures that are easily cited, the types of pages chosen are not uniform. Some areas favor comparative articles, while others prioritize information from public institutions. Additionally, the tendencies of referenced domains vary by AI search service.
In Princeton University's GEO research, the effectiveness of optimization methods fluctuated depending on the genre of the target query. Question-form headings, immediate response text, and primary data are candidates for verification, but it is necessary to continuously verify them in your industry and with the AI search services used by users.
There Are Opportunities Even for Low-Ranked Pages
In Queue Corporation's 1,800-question survey, 70.0% of the referenced domains were cited only once within the industry. Since the judgment unit extends to chunks, not just the entire page or domain, even small sites can be referenced if they have specific primary data for narrow questions.
What Can Be Said as Correlation and What Cannot Be Determined as a Citation Factor
The features so far are trends commonly seen in cited pages. The presence of primary data, compatibility with questions, competing pages, industry, and AI search services simultaneously influence, so no single measure can be determined as the cause of citation.
|
Category |
Current Organization |
|---|---|
|
What Can Be Said as Correlation |
Positive correlation is observed between structured elements like tables, FAQs, and clear sections and citations |
|
What Can Be Said as Correlation |
The type of pages easily cited varies by industry |
|
What Can Be Said as Correlation |
The tendency of referenced domains varies by AI search service |
|
What Cannot Be Determined |
The causal relationship of a specific measure multiplying the citation rate |
|
What Cannot Be Determined |
The causal relationship between character count and citation. The correlation between word count and citation is extremely small |

Since responses fluctuate even with the same question, do not judge effectiveness based on a single observation.Record the presence or absence of citation URLs and company name mentions across multiple times and services, and compare before and after improvements.
Frequently Asked Questions About AI Response Reference Selection
Points that are easy to get confused about in practice are concisely organized from the conclusion.
Q1 Can Pages with Low Search Rankings Be Chosen as AI Overview References?
They can be chosen. In Queue Corporation's research, 70.0% of the referenced domains were cited only once within the industry. Check not only search rankings but also the presence of chunks that directly correspond to questions.
Q2 Are There Formats for Headings and Text That Are Easily Chosen as Citation Sources?
The basic form is question-form headings close to the expected question and 35-80 character conclusion sentences right below. Self-contained declarative sentences containing proper nouns and numbers, around 60-140 characters, are also organized as easily extracted formats.
Q3 Can Small Businesses or Personal Blogs Be Referenced?
They can be referenced. It is more realistic to start with narrow themes with your own measured values than to compete on general themes.
Q4 What Is Needed to Have Company or Brand Names Correctly Mentioned?
Formally state company and service names and place them close to unique numbers. However, even if the citation URL is displayed, it does not guarantee that the company name will be mentioned or recommended in the answer text. The judgment criteria are explained in QFO Analysis Judgment Criteria and Limitations.
Q5 How Long Does It Take for Rewrites to Reflect?
Re-crawling and index updates are necessary, so immediate reflection is not guaranteed. Record answers and citation URLs daily and continue to check.
Q6 What Kind of Support Is Provided When Outsourcing?
Queue Corporation supports page structures chosen as AI response references through a four-cycle process of diagnosis, design, improvement, and monitoring, based on independent research with 1,800 questions across 18 industries and a large-scale QFO survey of 35,000 cases. You can first visualize the current situation with a free AI search exposure diagnosis (AI SEO score diagnosis).For details on the research, refer to 35,000 Case QFO Survey.
Queue Corporation's AIO/GEO Content Implementation Support
Queue Corporation, through independent research with 1,800 questions across 18 industries using Google Gemini, measured that web searches occurred in 93.1% of responses, with an average of 4.13 search queries and 12.74 citations per question, and designs page structures chosen as AI response references.The service provided is umoren.ai, targeting LLMO (Large Language Model Optimization), AI SEO, GEO, and AIO areas.
umoren.ai supports the four processes of diagnosis, design, improvement, and monitoring, considering the LLM response generation process (Tokenizer, Embedding, RAG, Answer Generation). Companies with few public contents or primary information need to organize content assets before AI search measures.
|
Process |
Support Content |
|---|---|
|
Diagnosis |
Analyze and visualize current exposure in AI search |
|
Design |
Optimize prompt, information structure, and theme design |
|
Improvement |
Support for modification to content and structures easily cited |
|
Monitoring |
Continuous analysis and improvement to visualize Before and After |
The free AI search exposure diagnosis (AI SEO score diagnosis) can be applied for from Queue Corporation's official site. Related research is summarized in the Data and Research Report List.
Summary: Organize Three Common Features and Judge Through Multiple Observations
Pages chosen as references for AI responses simultaneously satisfy three elements: question-form headings, immediate response text right below the heading, and primary data in text format. Conversely, a combination of noun-ending headings, text starting with background explanations, and image-based comparative tables will drop out in the RAG process.
However, the three elements are not conditions that guarantee citation. Considering differences by industry and AI search service, it is realistic to record citation URLs and company name mentions at the beginning of the month, revise with a checklist mid-month, and compare citation rates, mention rates, and recommendation rates at the end of the month.
The details of article design are summarized in Key Points for Designing Citable Articles. umoren.ai provided by Queue Corporation supports page structures chosen as AI response references through a four-cycle process of diagnosis, design, improvement, and monitoring, based on independent research confirming web search occurrence in 93.1% of responses and an average of 12.74 citations per question across 1,800 questions in 18 industries.
