Crawling and index control
Diagnose discovery, indexing, canonicalization, and preview-control problems before changing content.
Short definitions are useful. But when you're fixing rankings, you usually need the details. This wiki expands our glossary into practical, example-driven notes.
These paths connect definitions that belong to the same diagnostic or implementation decision. Use the glossary for a quick lookup; use the wiki when you need boundaries, examples, and validation steps.
Diagnose discovery, indexing, canonicalization, and preview-control problems before changing content.
Separate visibility, citations, referral traffic, and user outcomes instead of treating them as one metric.
Choose between documentation, APIs, skills, browser tools, MCP, and A2A based on the job and runtime.
An A2A Agent Card is a JSON document that describes an A2A server's identity, endpoint, capabilities, skills, interfaces, and authentication requirements.
Agent commerce signals are protocols and metadata patterns that help AI agents discover products, authenticate, pay, and complete transactions safely.
An Agent Skills index is a publisher-maintained catalog of real SKILL.md packages; the open specification defines each skill package, not one universal website discovery URL.
API Catalog SEO makes a genuinely public API understandable through stable documentation, OpenAPI, examples, access rules, versioning, and intentional discovery links.
Auth.md and OAuth metadata help AI agents understand authentication, authorization servers, protected resources, scopes, and consent boundaries.
DNS-AID is an informal label, not an adopted standard. Evaluate DNS-based agent discovery, its constraints, and safer established alternatives.
Link headers are HTTP response fields that help AI agents discover machine-readable resources such as sitemaps, llms.txt, API catalogs, and policies.
An MCP server card is an agent-readable summary of an MCP server's tools, resources, authentication requirements, and safe usage boundaries.
WebMCP is an experimental browser API for registering JavaScript tools that compatible AI agents can invoke in the context of an open web page.
Answer Engine Optimization is an informal practice, not one ranking system. Separate access, retrieval, answer accuracy, citations, and business outcomes with reproducible tests.
Agentic search is the shift from single-query search to AI agents that plan, search, read, and act on your behalf across many tools and sources.
AI brand mentions—being named in answers from ChatGPT, Perplexity, and AI Overviews—are now a measurable ranking and trust signal. Track and grow them.
AI citations are links or attributions shown with some generated answers. Verify what each citation supports, save reproducible tests, and separate visibility from referral outcomes.
AI providers may use separate agents for training, search indexing, and user-triggered fetches. Verify official documentation, robots rules, IP evidence, server logs, and outcomes.
Understand what Google confirms about AI Overviews, which SEO foundations still apply, which optimization myths to avoid, and how to measure visibility without promising citations.
Measure visits from AI-assisted products without assuming they are always identifiable or higher value. Audit referrers, direct traffic leakage, landing pages, and conversions.
Audit branded search results across markets and devices. Separate owned, earned, paid, and third-party results; correct identity conflicts and improve useful conversion paths.
ChatGPT search can cite public web pages. Learn the role of OAI-SearchBot, the limits of crawler access, and a defensible measurement workflow.
Citation rate measures how often a page or domain appears as a visible citation in a defined prompt set. It is a custom sampling metric, not an official Google or OpenAI ranking score.
ClaudeBot collects public web content for potential model development. Compare it with Claude-SearchBot and Claude-User before changing robots.txt.
Conversational search carries context across follow-up questions. Learn how to map user tasks, test changing intent, build useful paths, and avoid unsupported formatting hacks.
Measure LCP, INP, and CLS with field and lab data, diagnose page-level causes, and understand what Core Web Vitals can and cannot prove about Google rankings.
DeepSeek publishes open model weights and operates chat and API services. Learn what site owners can verify, what remains undocumented, and how to monitor visibility responsibly.
Digital PR earns brand mentions, authoritative links, and—increasingly—AI citations by getting your story into publications that matter.
Use Google's E-E-A-T guidance correctly: understand the role of experience, expertise, authoritativeness, trust, YMYL, quality raters, and practical evidence.
Embeddings are vector representations of data used to compare meaning and similarity. They matter in retrieval systems, but they are not a Google ranking score publishers can tune.
Entity SEO clarifies real people, organizations, products, and relationships with consistent facts, first-party pages, structured data, and verifiable external references.
GEO is an emerging label, not a universal ranking system. Trace the term to its research origin, separate access from citation, and measure AI visibility with reproducible tests.
GPTBot is OpenAI's model-training crawler, not its search crawler. Learn how to separate training, search visibility, and user-requested access.
Google no longer documents a standalone Helpful Content Update to optimize for. Audit affected pages, intent, originality, evidence, trust, and production quality instead.
Use hreflang to tell search engines which language and region a page targets. Avoid common mistakes that break multilingual SEO.
Information gain is the unique, additive value your content brings to a topic. Google explicitly rewards pages that contribute something the SERP does not already have.
The Knowledge Graph is Google's structured database of entities and their relationships. It powers knowledge panels, entity cards, and rich results.
A knowledge panel is the boxed information shown on the right of Google results for entities like brands, people, and places. Earn and shape it.
LLM hallucination is when a language model produces confident but false information. For SEO, it shapes how answer engines pick and trust sources.
LLM optimization is an informal practice, not one ranking system. Separate crawl access, retrieval, model output, citations, and conversions with reproducible tests and evidence.
llms.txt is a proposed file that gives LLM-driven crawlers a clean, machine-readable summary of your site—analogous to robots.txt for the AI era.
The Model Context Protocol (MCP) is an open standard that lets AI agents connect to tools, data sources, and apps in a consistent, secure way.
Multimodal search accepts or returns more than text. Use evidence-based image and video indexing practices without inventing universal ranking signals.
Learn how language models can support query and document understanding, what Google actually discloses, and why clear evidence matters more than speculative NLP optimization scores.
Negative SEO is the practice of attacking a competitor's rankings. Recognize common tactics and protect your site with practical defenses.
Understand rel=nofollow, sponsored, and ugc, choose the correct link qualification, and avoid treating nofollow as an indexing or PageRank control.
Off-page SEO includes backlinks, brand mentions, and entity associations. Learn what matters and what to ignore.
On-page SEO covers everything you control on a page: titles, headings, content, internal links, and structured data.
Organic search drives the most consistent, compounding traffic. Learn how it works and what it takes to win placements.
The People Also Ask box is a high-visibility SERP feature. Learn how it works and how to structure content to appear in it.
Page speed affects rankings, bounce rate, and revenue. Learn the metrics Google uses and the fixes that actually move the needle.
Pagination is how you split long content across pages. Implement it correctly with rel=next/prev, view-all, and canonical signals.
Parasite SEO is the practice of publishing content on high-authority third-party platforms to piggyback on their rankings. Risky, but still widely used.
Passage indexing lets Google rank individual passages within a page. It changes the unit of optimization from page to paragraph.
Understand Perplexity's web-backed answers, crawler controls, citations, and changing search modes. Test source visibility and referral outcomes without unsupported ranking claims.
PerplexityBot surfaces websites in Perplexity search and is separate from Perplexity-User. Learn the robots.txt, WAF, and log checks that matter.
Position zero refers to the SERP feature that appears above the first organic result—featured snippets, AI Overviews, PAA, and more.
Prompt engineering is the practice of crafting inputs that steer large language models toward accurate, useful, and safe outputs. It is a core SEO-adjacent skill.
Quality Score estimates ad relevance and landing page experience. While it lives in Ads, the same signals improve organic SEO too.
Query decomposition means breaking one complex question into smaller answerable parts. Use it to improve one strong page, not to justify thin subquery pages.
Query understanding is how search systems interpret intent, entities, language, and context. Use it to improve page clarity, not to chase an imaginary score.
A search query is what a user types into a search engine. Understanding query intent is the starting point for any SEO strategy.
Learn how to find question-based keywords from Search Console, validate the SERP, cluster related questions, and map them to pages without stuffing FAQ blocks.
Ranking factors are the signals search engines use to order results. Learn the ones that matter and how to weigh them.
Rich snippets add stars, prices, FAQs, and more to your listing. Implement them with structured data to lift CTR and visibility.
Robots.txt is a file that tells crawlers which paths they may request. Use it carefully—a wrong rule can deindex your site.
Schema markup is structured data that helps search engines understand your content. Use it to unlock rich results and knowledge panels.
Search Everywhere Optimization is the practice of optimizing for every surface where people search—Google, YouTube, TikTok, Amazon, ChatGPT, Perplexity, and beyond.
SGE was Google's experimental generative Search experience. Learn how it evolved into AI Overviews and AI Mode, and what site owners can verify today.
Semantic SEO is an informal approach to clarifying topic, intent, entities, and relationships. Use it to improve page clarity, not to chase a semantic score.
The SERP is the page that shows search results. Modern SERPs include featured snippets, PAA, image packs, and AI overviews.
Share of Model is a metric that measures how often your brand appears in AI-generated answers for a given topic set, similar to Share of Voice for traditional search.
Site reputation abuse concerns third-party content published mainly to exploit a host site's established ranking signals. Audit ownership, purpose, promotion, oversight, and search intent.
Source-of-truth content is structured, authoritative content that AI engines repeatedly cite. It is the new foundation of brand visibility in AI answers.
An SSL certificate enables HTTPS, which is a confirmed Google ranking signal and a baseline trust signal for users.
Structured Data is standardized markup that helps search engines and AI systems understand the meaning of your pages, not just their words.
Technical SEO covers crawling, indexing, rendering, site speed, and structured data. Get it right and the rest of SEO gets easier.
Thin content is low-value, shallow, or auto-generated content. Identify and improve or remove it to lift overall site quality.
The title tag is the clickable headline in search results. Write one that matches intent, includes the topic, and fits the pixel limit.
A topic cluster is a pillar page supported by a group of interlinked articles on subtopics. It is the classic way to build topical authority.
Topical authority is the perception—shared by users and search engines—that your site is the definitive source on a given subject. Build it with depth and structure.
A clean URL structure helps users and search engines. Use short, descriptive slugs and avoid parameter soup.
User experience drives engagement, conversions, and rankings. Learn the Core Web Vitals and UX signals Google measures.
User intent is the goal behind a query: informational, navigational, commercial, or transactional. Match it to rank and convert.
Vector search retrieves similar items by comparing embeddings. It helps explain modern retrieval systems, but it is not a Google ranking score publishers can tune.
Vertical search engines focus on a specific content type: images, video, news, shopping, or local. Optimize for the ones your audience uses.
Visual search lets users search with images instead of text. Optimize images with alt text, structured data, and high-quality visuals.
Voice search queries are longer, more conversational, and question-based. Optimize for natural language and direct answers.
A web crawler is a bot that discovers and indexes web pages. Understand how it works to make your site easy to crawl.
Web standards are the official guidelines (HTML, CSS, accessibility) that keep the web consistent. Following them is good SEO and good citizenship.
White hat SEO follows search engine guidelines and focuses on users. It builds durable rankings without the risk of penalties.
The X-Robots-Tag HTTP header controls indexing and snippetting without HTML. Use it for non-HTML files and edge-case crawling rules.
XHTML is HTML written in XML syntax. Strict, but largely replaced by HTML5. Learn when it still matters and when to use HTML5 instead.
An XML sitemap lists the URLs you want crawled. Submit it in Search Console to help search engines discover and prioritize your content.
Yahoo Search is powered by Microsoft Bing. Most SEO work transfers, but Yahoo has its own portal audience worth considering.
Yandex is a Russian search engine with its own Webmaster workflows. If you target Russian-speaking users, verify robots.txt, noindex, canonical, localized pages, and reindexing in Yandex Webmaster.
YMYL pages can impact a person's wellbeing, so Google holds them to a higher E-E-A-T standard. Identify yours and raise the quality bar.
A zero-click search is one where the user gets the answer on the SERP. Win them with structured data, PAA, and direct, scannable answers.
A zero-result SERP shows no organic results. Google provides a direct answer for queries it can answer with high confidence.
A zone file lists the DNS records for a domain. It is the source of truth for how a domain name resolves to services on the web.
AEO is the practice of optimizing content to appear as direct answers in search engines, featured snippets, voice search, and AI chatbots.
GEO is the practice of optimizing content to appear in AI-generated search responses from Google SGE, ChatGPT, Perplexity, and other generative engines.
Image optimization reduces file size while preserving visual quality. Learn compression formats, responsive images, lazy loading, and Core Web Vitals impact.
Junk content is not a formal Google policy label. Learn which official spam and quality signals it usually points to, how to audit low-value pages, and when to rewrite, merge, or keep them.
Plan link acquisition around useful assets, relevant audiences, editorial independence, commercial qualification, risk review, and outcomes beyond raw backlink counts.
Build a defensible local SEO program around Google Business Profile accuracy, relevance, distance, prominence, reviews, landing pages, and measured customer actions.
Long-tail keywords are longer, more specific search phrases with lower competition and higher conversion rates. Learn how to find and target them effectively.
Meta descriptions don't directly affect rankings but heavily influence click-through rates. Learn how to write compelling, accurate descriptions that convert.
Meta tags provide search engines with information about your pages. Learn which tags matter, which don't, and how to implement them correctly.
Google uses the mobile version of your site for indexing and ranking. Learn how to ensure your mobile experience delivers equal content, structured data, and metadata.
Query fan-out is Google's technique for expanding one search into multiple related retrieval paths. Plan one strong page around the real task, not one page per variation.
Learn how Retrieval-Augmented Generation works in Google Search and why it matters for your content's visibility in AI-powered search results.
Learn what scaled content abuse is, how Google detects it, and how to avoid violating this critical spam policy in your SEO strategy.
AI SEO can mean using AI in SEO workflows or improving visibility in AI-assisted search. Define the objective, validate evidence, govern automation, and measure real outcomes.
Learn what GA4 engagement rate measures, how it differs from bounce rate, and how to use it for page diagnosis without treating it as a Google ranking factor.
Use outbound links to support claims and help readers, choose sponsored, ugc, or nofollow correctly, and audit destination status, context, accessibility, and disclosure.
Featured snippets are the highlighted answers that appear at the top of Google search results. Learn how to optimize your content to earn this highly visible position.
Learn what a website footer should do, how the HTML footer element works, when footer links help navigation, and how to audit footer SEO without stuffing links.
Freshness is a ranking factor that considers how recent your content is. Learn how to optimize your content for freshness and when it matters most.
Gateway pages are low-quality pages created solely for search engines. Learn why they're considered spam and how to avoid penalties.
Geo targeting helps you reach users in specific locations. Learn how to optimize your website for local search and improve your local SEO rankings.
Learn how to optimize your content for large language model search and generative AI platforms to improve visibility in AI-powered search results.
Learn how to optimize your website for multiple languages and regions, reaching global audiences while maintaining strong search rankings.
Dofollow isn't an official attribute. Learn how link equity works, what nofollow really does today, and how to audit your links.
Duplicate content confuses indexing and can split ranking signals. Learn common duplicate patterns and practical fixes using canonical, redirects, and internal linking.
Domain Authority (DA) is a Moz metric, not a Google ranking factor. Learn how to interpret DA/DR-style scores without fooling yourself.
CTR is the percentage of impressions that become clicks. Learn what affects CTR in Google Search and how to improve titles, snippets, and intent match.
Crawling is how bots discover your pages. Learn crawl paths, crawl budget basics, and common blockers like robots.txt, noindex, and soft 404s.
Breadcrumbs help users and crawlers understand site structure. Learn when to use breadcrumbs, how to mark them up, and common implementation mistakes.
Learn what bounce rate means in GA4, why a one-page visit is not automatically a bounce, and how to diagnose high bounce without treating it as a Google ranking factor.
Backlinks still matter, but not all links are equal. Learn what makes a backlink helpful, how anchor text plays in, and red flags to avoid.
Anchor text is the clickable text in a link. Learn practical anchor text patterns for internal and external links, plus common mistakes.
In SEO, an algorithm is the system that decides what ranks. Learn how to think about algorithms, signals, and updates without chasing myths.
Learn how Googlebot crawls sites, what limits and priorities it uses, and how to check crawl issues without guesswork.
Use H1-H6 headings to label sections, provide an accessible outline, and clarify page content. Audit rendered headings without treating hierarchy as a ranking trick.
Understand how Google handles 2xx, redirects, 4xx, 429, and 5xx responses, then diagnose soft 404s, redirect chains, crawl slowdowns, and indexing loss.
Separate discovery, crawling, indexing, and serving. Diagnose noindex, robots.txt, canonical, rendering, duplicate, and quality-related exclusions.
Build crawlable internal links with descriptive anchors, useful context, direct destinations, and measurable site architecture. Diagnose orphan pages, weak hubs, loops, and broken paths.
Understand Google crawling, rendering, and indexing for JavaScript sites. Test raw and rendered HTML, status codes, links, canonicals, blocked resources, and hydration failures.
Learn how to implement JSON-LD, choose eligible Schema.org types, connect entities, validate rich-result requirements, and monitor structured data safely.
Keyword density describes term frequency but has no documented ideal ranking percentage. Review repetition, context, clarity, and user value instead of chasing a tool score.
Keyword stuffing fills pages with terms or numbers to manipulate rankings. Learn Google's policy examples and audit titles, copy, metadata, links, templates, and hidden text.
Keywords are the bridge between your content and what searchers are looking for. Learn how to think about them without obsessing.
Alt text helps search engines and screen readers understand images. Learn how to write useful, non-spammy alt text with examples and quick audits.
We use cookies to enhance your experience. By continuing to visit this site you agree to our use of cookies.