On this pageTable of contents+
- Day 1: Crawling
- Search Has Reached an Inflection Point
- AI Overviews Are Not a Separate Search System
- Human Writing Remains the Basis of Ranking
- Multimodal Content Is Becoming a Baseline
- Gen Z Is Moving Beyond the Search Box
- Search Console's Data Cadence and Role
- How Crawl Budget Was Described
- Longer Query Structures Are Growing
- Google Said It Was Not Adopting LLMs.txt
- AI Is Not an Exception to the Search Pipeline
- Day 2: Indexing
- How Google Understands Web Content
- The Position of Main Content Can Affect Its Interpretation
- Deduplication and Canonical Priorities
- How Was a Soft 404 Defined?
- Ways to Control Indexing and Their Limits
- Use Structured Data, but Use It Carefully
- How Are Images and Videos Indexed?
- How Geotargeting Can Affect Content Selection
- The Signal-Extraction Stage
- What Kind of Content May Not Be Indexed?
- Is E-E-A-T a Metric?
- The Google Trends API Alpha Was Announced
- Day 3: Serving
- How Queries Are Understood
- Understanding the Quality Signal
- Four Pillars of Quality
- E-E-A-T and the Priority of Trust
- Three Reasons Google Search Systems Are Updated
- How to Respond After an Update
- The Boundaries and Misconceptions of Structured Data
- Overall Summary

In the week before the source article was published, Google held the three-day Search Central Live Deep Dive APAC 2025 event in Bangkok. Its themes were crawling, indexing and serving, and the sessions covered both underlying systems and practical details.
We prepared this long-form summary because, in SEO, written material can still be the best format for slowing down and thinking carefully.
The more fragmented the information becomes, the more it needs to be organized systematically. The faster the field moves, the more important it is to keep learning.
We should also add one caution: a statement from Google should not automatically be treated as infallible. SEO rarely offers absolute answers, so practitioners need to develop their own judgment and principles.
This dense summary took substantial time and effort to compile. If you find it useful, please share it with others who may benefit.
With that context in place, here are the notes.
Day 1: Crawling
Keywords: infrastructure for AI search experiences, changing Gen Z entry points, structured multimodal content, crawl-budget logic and the role of Search Console.
One-sentence summary:
AI is changing where some search journeys begin, but the conference notes emphasized that Google's core processes for crawling, understanding and ranking have not been rewritten. High-quality content created for people remains the starting point.
One often-overlooked point:
Frequent AI-powered search experiences do not mean that publishers should write for an AI system. The session notes said that ranking systems are not trained on AI-generated text and reiterated that the intended reader should be a person, not a model. Treat this as a report of the session rather than a complete public specification of Google's training data or ranking systems.
Search Has Reached an Inflection Point
Mike Jittivanich described search as being at the intersection of three forces:
1. AI innovation with an impact comparable to the shifts to mobile and social platforms. 2. Users expecting faster, more natural and more conversational search experiences. 3. Gen Z search behavior differing markedly from that of previous generations.
Users are no longer limited to conventional blue links. The challenge described for search teams was how to combine AI experiences, image retrieval and other new formats with blue links in one interface without creating conflict.

AI Overviews Are Not a Separate Search System
- The session said that Googlebot crawls content used for both traditional blue links and AI search features.
- It described AI Overviews, AI Mode and blue links as sharing core crawling, indexing and serving infrastructure.
- Other crawlers in Google's wider ecosystem were described as serving Gemini and large language model use cases.
- The notes said the common pipeline parses and renders HTML, removes duplicates, processes ranking signals and detects spam.
The presentation also described Google as bringing DeepMind reasoning models into Search to support multimodal, multi-step and task-oriented experiences, including shopping, dining plans and route planning in AI Mode.

Human Writing Remains the Basis of Ranking
• The session notes said ranking systems are not trained on AI-generated content. • They said AI-generated text can be indexed but is not used as ranking-model training material. • They characterized high-quality human-created pages as the basis for ranking models, highlighting natural language, clear structure and reliable information. These are attributed conference claims, not a complete description of Google's internal training processes.
Multimodal Content Is Becoming a Baseline
• Give images descriptive alt text. • Provide captions or transcripts for video where appropriate. • Use natural language that can support conversational and voice queries. • As search touchpoints fragment, make useful information understandable through visual, audio and semantic entry points.
Gen Z Is Moving Beyond the Search Box
• Slides cited in the notes reported 65% year-over-year growth for Google Lens searches, more than 100 billion searches during 2025 and commercial intent in about 20% of them. • They also reported that roughly 10% of Gen Z search journeys began with Circle to Search or another AI entry point. • The broader takeaway was that discovery no longer begins only in a text box, so content may need to support image-first, voice-first and other nonlinear journeys. The figures here are preserved as figures shown at the event and were not independently reproduced for this article.

Search Console's Data Cadence and Role
• The notes described a delay of about two days, based on Pacific Time. • They said near-real-time and finalized data could differ by around one percent. • They characterized Recommendations as modular, actionable guidance intended to help less experienced users. • The product cycle was summarized as user need, data preparation, design and development, testing, and launch. • Search Console was described as a bridge between Google Search infrastructure and a website that can be used to:
• Diagnose crawling issues. • Track search performance. • Investigate whether a page appears eligible for AI-powered search experiences. Search Console does not guarantee inclusion in any feature.
How Crawl Budget Was Described
• The session summarized crawl budget as the interaction of a crawl-rate limit and crawl demand. • The source notes claimed that only HTTP 5xx responses truly consume crawl budget and that 4xx responses do not, although they can affect scheduling priority. Current public Google documentation does not present that as an absolute rule, so use Crawl Stats and the current status-code guidance when diagnosing a site. • Broken links and slow responses were said to reduce crawling efficiency. • AI features may be associated with more crawling, but a higher crawl rate does not imply better rankings.
Longer Query Structures Are Growing
• The presentation reportedly said natural-language queries of more than five words were growing 1.5 times as quickly as shorter queries. • Users increasingly express intent as questions and tasks. • Content structures should therefore support questions, scenarios and task-oriented explanations. The growth figure is retained as a conference statistic.
Google Said It Was Not Adopting LLMs.txt
• The source incorrectly called llms.txt an IETF draft; it is a third-party proposal rather than an IETF standard. The substantive conference point was that Google was not participating in or planning to adopt it. • Google's public guidance continues to document robots.txt for controlling crawler access. • Not every AI crawler necessarily follows robots.txt, so publishers should review each operator's documented controls.
AI Is Not an Exception to the Search Pipeline
• The notes said Google handles pages used in AI search by parsing and rendering HTML, removing duplicates, analyzing language, applying reasoning models and detecting spam. • They said the wider Search system still interprets queries and scores content. • The visible difference may be the final format: a blue link, an AI Overview or an AI Mode module. These descriptions summarize the presentation and should not be read as a complete internal architecture.
Day 2: Indexing
Keywords: indexing, main-content recognition, duplicate-content clustering, crawler controls, signal extraction, structured data, SpamBrain, soft 404s and the Google Trends API.
One-sentence summary: the notes framed indexing as a judgment about whether content appears useful to people, not merely whether a technical mechanism made it discoverable.
One often-overlooked reminder: adding rel=canonical does not force Google to select that page. The speaker said Google considers hijacking risk and the experience for users before it weighs a site owner's signals. Public documentation describes canonical annotations as signals rather than commands.
How Google Understands Web Content
- HTML is normalized into a DOM, after which navigation, headings and main-content areas can be identified.
- Main content can include text, images, video or a functional module such as a calculator.
- JavaScript-rendered output can be processed before an indexing decision is made.
- The system extracts structural information such as
rel=canonical,hreflangandmeta robots. - Core page content can be separated from styling and navigation for later processing.
- The speakers continued to emphasize intuitive content structure and said responsive and adaptive layouts were not inherently favored over one another.
- The notes said Google tokenizes HTML content to build its index and traced that tokenization work to the Tokyo office in 2001, adding that related technology is also used in AI products.
The Position of Main Content Can Affect Its Interpretation
The speaker gave an example in which moving the term Hugo 7
from a sidebar into the main text coincided with a clear rise in overall visibility. The lesson offered was that terms in the main content, especially headings, subheadings and opening paragraphs, may be interpreted differently from sidebar material. Because this was an example rather than a controlled causal study, page structure and other changes should still be checked.
Deduplication and Canonical Priorities
The notes discussed redirects, content similarity and rel=canonical as inputs to duplicate-content handling. They said Google first considers hijacking risk, then user experience and finally the canonical preference supplied by a site owner. Even when a canonical is declared, Google may select another version. The source also warned that soft-error or template-generated pages could affect how a duplicate cluster is assessed. Public guidance describes redirects and rel=canonical as strong signals, not guarantees.

How Was a Soft 404 Defined?
If a page appears to exist but its main content is extremely thin or unhelpful, Google may treat it as a soft 404. The session used the term centerpiece annotation to describe a problem associated with the main content rather than an HTTP error. The source further said a soft-404 assessment could affect a duplicate cluster and that usefulness and trust both matter. These are the source's descriptions of the presentation; individual URLs should be diagnosed with Search Console and Google's current soft-404 documentation.
Ways to Control Indexing and Their Limits
• robots.txt controls whether a compliant crawler may request content, but it is not a dependable method for keeping a web-page URL out of search results. • A robots meta tag or HTTP header can control whether crawled content may be indexed, with directives such as noindex and noimageindex. • none is equivalent to noindex, nofollow. • unavailable_after can make time-limited or expired material ineligible after a specified date. • The source associated notranslate with suppressing translation prompts; Google's documented meta tag specifically prevents Google from offering a translated search-result version.
Use Structured Data, but Use It Carefully
Structured data provides explicit clues about a page's meaning and entities; it is not a direct ranking shortcut. Appropriate Schema markup can make a page eligible for supported search features, but redundant markup adds code without extra ranking weight. Structured data needs maintenance because stale, inaccurate or invalid values can remove eligibility. Even correct markup never guarantees a rich result: Google chooses whether and how to show one in context.
How Are Images and Videos Indexed?
The speaker said indexing begins with HTML and that a separate asynchronous pipeline handles media. If HTML has been indexed but an image or video has not appeared in Search, media processing may still be pending, although crawlability, indexability and technical errors should also be checked. The speaker reportedly said Google does not evaluate an image simply by whether it was AI-generated, but by whether it communicates useful content.

• The example in the notes was an AI-generated image used as a decorative prop on the first day. Although some details were inaccurate, the presentation said that decorative error did not affect ranking.
How Geotargeting Can Affect Content Selection
• The session described hreflang as a central bridge between regional versions. • It said equivalent regional pages do not need to be hidden merely to avoid duplication because the system can recognize the configuration. • However, two pages that change only the domain and provide no localization can confuse both systems and users. • Geographic targeting and language targeting are different and should not be treated as substitutes. • The presentation listed several possible regional signals:
• A country-code top-level domain such as .sg or .au. • hreflang annotations. • Server location. • Page language, currency, Business Profile information and the regional context of links. The speaker did not quantify how Google currently weighs or uses these signals, so implement them according to Google's current guidance for international sites.

The Signal-Extraction Stage
After structural processing, the system was described as extracting direct and indirect signals, including:
• PageRank, which the notes said remains in use internally. • Terms on the page and their semantic positions. • Links and mentions.
The session said the signal called PageRank is no longer identical to the original 1996 algorithm but still participates in ranking under the same name.

What Kind of Content May Not Be Indexed?
- A page carrying an effective
noindexdirective. - Expired event content, for which the source suggested
unavailable_afteras one management option. - Extremely thin or low-value content that is treated as a soft 404.
- A duplicate that is not selected as the canonical representative of its cluster.
- Content that SpamBrain identifies as spam; a slide reportedly cited 40 billion spam pages detected per day. That figure is retained as a conference-slide claim rather than an independently verified current total.

Pages that violate policies, such as pages offering malicious downloads or using misleading titles, were also listed.
The source added that internal links help only when a page already provides some value. A link alone does not preserve indexation; it can help discovery and context, but the destination still needs to be useful.
Is E-E-A-T a Metric?
No. The Search Quality Rater Guidelines use E-E-A-T to ask whether content shows firsthand use and expert knowledge, comes from a recognized source and can be trusted. It is not one indexing or ranking parameter. Google also states that E-E-A-T is not one specific ranking factor and that rater data is not used directly in ranking algorithms.
The Google Trends API Alpha Was Announced
Daniel Waisberg and Hadas Jacobi announced the alpha:
- It uses consistently scaled search-interest data so separate requests can be compared without each one being rescaled independently.
- It provides a rolling five-year window with data available through approximately two days ago.
- The source listed weekly, monthly and yearly aggregation; Google's announcement also lists daily aggregation.
- It supports analysis by region and subregion.
- It creates a programmatic path for trend monitoring and comparisons over time, but access began as a limited alpha rather than a generally available service.
Day 3: Serving
Keywords: query understanding, contextual synonyms, the E-E-A-T quality framework, types of search update and the limits of structured data.
One-sentence summary: Google Search uses many signals, and the sessions presented quality as one of the most important. The central question is whether users have reason to trust the content.
One often-overlooked point: Google may show information such as a site name or breadcrumbs even without structured data. Clear content and sound page structure can matter more to understanding than simply adding Schema markup.
How Queries Are Understood
The notes divided Google's serving process into five stages: query understanding, retrieval, index selection, ranking and application of search features such as rich results.

Query understanding begins with segmentation. For languages such as Chinese and Japanese that do not use spaces between every word, the notes said Google learns segmentation from historical queries and documents; not every language requires the same step. The system can then remove stop words unless they form part of a specific phrase or entity, such as The Lord of the Rings. Query expansion can introduce cross-language synonyms to better interpret intent. One mechanism described was contextual synonyms: car hire
and rental car
may not be identical dictionary entries, but observed behavior can identify them as siblings that are interchangeable in a particular context.


Users do not see this language analysis and semantic mapping, but it can improve the retrieval of relevant information.
Understanding the Quality Signal
The speaker said many signals contribute to ranking and emphasized the quality of a page. The notes grouped good content into five areas:
- People-first purpose.
- Expertise.
- Content and quality.
- Presentation and production.
- Avoiding search-engine-first content.
The notes connected these dimensions to the Search Quality Rater Guidelines. The guidelines do not directly calculate rankings; they help raters assess whether systems are producing good results. Updates to the guidelines can reflect an evolving description of what good content looks like.

Four Pillars of Quality
The source summarized four quality dimensions discussed in the guidelines:
- Effort: evidence of the time, skill and experience invested in creating the content.
- Originality: independent research, original analysis or a distinctive perspective.
- Talent or skill: direct experience and capable writing, even when the creator is not a formally recognized expert.
- Accuracy: claims grounded in evidence and consistent with relevant expert or public consensus.
E-E-A-T and the Priority of Trust
Google's public E-E-A-T guidance identifies trust as the most important aspect. The term also covers experience, expertise and authoritativeness. The source notes extended the point beyond YMYL topics and said content that clearly contradicts established expert consensus may be judged unreliable. E-E-A-T remains a framework, not one measurable ranking score.
The session also clarified that a site having many 404 responses or using noindex does not by itself define content quality. Those are technical states; their operational effects and the quality of the remaining content still need to be evaluated separately.
Three Reasons Google Search Systems Are Updated
The source grouped Search updates into three broad purposes:
1. Supporting new content formats: when demand grows for formats such as short video or interactive visuals, Google may add new search experiences. 2. Improving relevance: core updates can change how systems identify useful content across Search rather than targeting one particular site. 3. Reducing low-quality content and spam: systems may be adjusted to identify and demote content that provides little value or attempts to manipulate results.

How to Respond After an Update
A core update is not necessarily a penalty, and the source cautioned against thinking of improvement as a simple return to an earlier ranking state. It summarized Google's advice as follows:
- Continue creating high-quality content.
- Study pages that serve users better and understand what they do well.
- For spam updates, first identify the relevant policy or documented issue.
- Correct or remove material that violates the applicable official guidance.
The Boundaries and Misconceptions of Structured Data
- Adding structured data does not directly guarantee higher rankings.
- When supported and appropriate, it can make a page eligible for enhanced search presentations that may affect how users interact with the result.
- Schema markup does not guarantee a rich result; the search system decides whether to show one.
Key reminders:
- Even without structured data, Google may infer and display elements such as a site name or breadcrumbs from page and site signals.
- Valid HTML and clear content structure remain fundamental to understanding.
- Schema is not a one-time task. Review it regularly so that visible content and markup remain accurate, consistent and eligible for supported features.
Overall Summary
Across all three days, the discussions repeatedly returned to the same question:
When the way content is presented changes dramatically, how does Google decide whether that content deserves to be seen?
Whether the topic is AI Overviews, structured data, E-E-A-T or multimodal retrieval, the source distilled the discussion into two questions:
• Was the content written for people? • Does the content deserve their trust?
This is not the end of SEO. It is a return to what SEO should have been about from the beginning.
