How the homepage tells a crawler who it is

Short answer: The homepage was already allowed in search. What a crawler could not do was quote it, or find the map of the site. /robots.txt and /sitemap.xml never reached WordPress, so both were Apache 404s. A real robots.txt now points at the sitemap WordPress already knew how to serve. The homepage also carries one description, in the meta tag, in the link preview, and in schema.org data. A plain llms.txt says the same thing for an AI crawler. None of this changes the layout, and none of it forces a recommendation.


What a crawler already had

Search was not discouraged. The WordPress setting blog_public is on. The title was already HOHOLIVE – 好好生活. The canonical URL was already https://hoho.live/. The page language is en-US, with traditional Chinese in the welcome line and the site name.

What was missing was a description. Without one, a search result and an AI answer have to guess from the title, or from whatever sentence they happen to lift. The homepage had no meta name="description", no Open Graph tags, and no structured data.


The file crawlers ask for was a 404

Pretty permalinks are off, so pages are ?page_id=. WordPress can still build a sitemap, but only on the query form https://hoho.live/?sitemap=index. The same is true of robots: /?robots=1 answers, /robots.txt does not.

A request for /robots.txt never arrives at WordPress. Apache answers with its own 404. The same happens for /sitemap.xml and for any pretty path such as /privacy-policy/. Turning pretty permalinks on, without teaching the server to hand those paths to index.php, would make the real pages 404 as well. So the permalink structure stays plain.

Crawlers do not try ?robots=1. They fetch /robots.txt. That URL now has a file of its own, in the site root:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://hoho.live/?sitemap=index

Nothing is blocked except the admin screens. AI crawlers are not singled out and not refused. The sitemap line is the query URL that already works. It lists the homepage, the privacy policy, and two other public pages.


One description, three places

The sentence is written once, in the must-use plugin that already holds the homepage decoration, and printed three ways.

The meta description is what a search snippet can use:

<meta name="description" content="Linux and macOS administration, automation, and infrastructure. HoHo's personal site, 好好生活. Focus on what I know, what I can do and what I have.">

Open Graph and a Twitter summary card repeat it, so a chat app or a link preview does not have to invent a blurb. The image is the site icon that was already there, not a new picture.

The third copy is JSON-LD, schema.org, in the homepage head. Three objects, tied together by id:

  • A WebSite named HOHOLIVE, alternate name 好好生活, languages English and traditional Chinese.
  • A Person named HoHo, with the email already printed on the page, and a short list of what the page itself says he works on: Linux, macOS, infrastructure, automation.
  • A WebPage for https://hoho.live/, part of that site, about that person.

The two paintings had empty alt. They now say what is in the picture: birds on a flowering branch, and a child playing a flute among bamboo. The files were not edited. Only the text a crawler can read.


A plain text file for an AI crawler

Some AI crawlers look for /llms.txt, a short Markdown file, before they spend a fetch on the HTML. It is a convention, not a ranking switch. hoho.live/llms.txt states the same facts as the description: who the site is, the focus line, the contact address, the public pages, and that LUDA is run by the same person. Mail and Admin are left out. They are not pages to recommend.


What this does not do

A crawler can now tell that this is HoHo’s personal site, what the work is, and which URLs are public. That is the part we can hand over. Whether a search engine shows the page, or an AI quotes it, is still that engine’s choice. The headings on the page were left as they are, so the column of text did not change shape. Pretty URLs wait until the server actually routes them to WordPress.

Leave a Reply

Your email address will not be published. Required fields are marked *