llms.txt generator
Enter your website address: the tool reads your robots.txt, your sitemap and up to 30 pages (title, meta description), then writes an llms.txt in the llmstxt.org format and a robots.txt block for AI crawlers. Free, no sign-up, nothing is stored.
Example: https://llmstxt.org — only public sites are read; our crawler follows your robots.txt.
Your llms.txt
Publish it at the root of your site: https://yoursite/llms.txt (plain text, UTF-8).
robots.txt block for AI crawlers
Report
Current llms.txt
Pages without their own meta description
llms.txt then uses the first paragraph. A meta description (about 150 characters) improves both llms.txt and your search snippets.
Pages with errors or excluded
AI crawler access today (home page, current robots.txt)
URLs skipped before reading
How the generator works
- robots.txt is read first. Our crawler (
llms-txt-generator) skips every URL your rules disallow, and stops if robots.txt answers with a server error, as RFC 9309 requires. - Sitemap: the sitemaps declared in robots.txt, otherwise
/sitemap.xml. Sitemap indexes and.xml.gzfiles are supported. No sitemap? Internal links from the home page are followed. - Up to 30 pages are read: the home page first, then pages from every section of the site. For each page:
<title>, meta description (orog:description, or the first paragraph),noindex, canonical. - llms.txt is written in the llmstxt.org format: one H1 with the site name, a blockquote summary, H2 sections built from your URL folders (
/blog/→## Blog), legal pages in anOptionalsection. The output is checked against the format before you see it. - robots.txt block: 20 official AI crawler tokens (OpenAI, Anthropic, Perplexity, Google, Apple, Meta, Amazon, Mistral, Common Crawl) with three policies. Your current
User-agent: *rules are copied for allowed crawlers, because a crawler with its own group stops reading*.
llms.txt example
What the generator produces for a small shop site with a blog (shortened):
# Atelier Grès
> Handmade stoneware tableware from Normandy.
## Main pages
- [Atelier Grès](https://shop.example.com/): Handmade stoneware tableware from Normandy.
- [Contact](https://shop.example.com/contact): Opening hours, address and phone.
## Blog
- [Choosing a mug](https://shop.example.com/blog/choosing-a-mug): Capacity, handle, glaze: our criteria.
- [Caring for stoneware](https://shop.example.com/blog/stoneware-care): Hot water, no abrasive products, air dry.
## Optional
- [Legal notice](https://shop.example.com/legal): Publisher, host, contact.
More detail on the format, and on who actually reads it, in What is llms.txt?
FAQ
Will llms.txt improve my rankings or get me cited by ChatGPT?
Nobody can promise that. Google Search says it does not use llms.txt, and no AI company has committed to reading it. It is cheap to publish and may help agents and coding tools; the robots.txt rules are the part with a documented effect. Details and sources: What is llms.txt?
Is anything stored?
No. Your URL is used for one generation and the result is sent back to your browser. No database, no cookies, no analytics. Against abuse, a hashed (never clear) form of your IP is counted for at most one hour. Cloudflare, the host, keeps standard technical logs. Details.
Why 30 pages?
The generator runs in a free, short-lived serverless function with a budget of 45 requests and 25 seconds. 30 pages covers most small sites. For larger sites the result is a sample of every section, which you should edit.
Which sites are refused?
Private, local and reserved addresses (for example 127.0.0.1, 10.x, 192.168.x, 169.254.169.254, [::1]), host names like localhost, non-standard ports and anything other than http/https. This protects the service against server-side request forgery.
My site is a single-page app. Will it work?
Only if titles and descriptions are in the HTML the server sends. The generator does not run JavaScript.
How do I block our crawler?
Add User-agent: llms-txt-generator / Disallow: / to your robots.txt. See the crawler details.
Comment fonctionne le générateur
- robots.txt est lu en premier. Notre robot (
llms-txt-generator) ignore toute URL que vos règles interdisent, et s'arrête si robots.txt répond par une erreur serveur, comme l'exige la RFC 9309. - Sitemap : ceux déclarés dans robots.txt, sinon
/sitemap.xml. Index de sitemaps et fichiers.xml.gzacceptés. Pas de sitemap ? Les liens internes de l'accueil sont suivis. - 30 pages au plus : l'accueil d'abord, puis des pages de chaque rubrique. Pour chacune :
<title>, meta description (ouog:description, ou premier paragraphe),noindex, canonical. - llms.txt au format llmstxt.org : un H1 avec le nom du site, un résumé en citation, des sections H2 tirées de vos dossiers d'URL (
/blog/→## Blog), les pages légales dans une sectionOptional. Le résultat est contrôlé avant affichage. - Bloc robots.txt : 20 jetons officiels de robots IA (OpenAI, Anthropic, Perplexity, Google, Apple, Meta, Amazon, Mistral, Common Crawl) et trois politiques. Vos règles
User-agent: *sont recopiées pour les robots autorisés, car un robot qui a son propre groupe ne lit plus*.
Exemple de llms.txt
Voir l'exemple ci-dessus en anglais ; le générateur nomme la section principale « Pages principales » quand votre site est en français. Plus de détails dans Qu'est-ce que llms.txt ?
Questions fréquentes
llms.txt améliore-t-il mon référencement ou mes citations par ChatGPT ?
Personne ne peut le promettre. Google Search dit ne pas utiliser llms.txt et aucun éditeur d'IA ne s'est engagé à le lire. Le fichier coûte peu et peut aider des agents et outils de code ; les règles robots.txt, elles, ont un effet documenté.
Des données sont-elles enregistrées ?
Non. L'URL sert à une génération et le résultat est renvoyé à votre navigateur. Ni base de données, ni cookie, ni mesure d'audience. Contre les abus, une empreinte de votre IP (jamais en clair) est comptée une heure au plus. Cloudflare, l'hébergeur, conserve des journaux techniques. Détails.
Pourquoi 30 pages ?
Le générateur tourne dans une fonction serverless gratuite limitée à 45 requêtes et 25 secondes. Pour un grand site, le résultat est un échantillon de chaque rubrique, à relire.
Quels sites sont refusés ?
Les adresses privées, locales et réservées (127.0.0.1, 10.x, 192.168.x, 169.254.169.254, [::1]…), les noms comme localhost, les ports non standard et tout ce qui n'est pas http/https : protection contre la falsification de requêtes côté serveur (SSRF).
Comment bloquer votre robot ?
Ajoutez User-agent: llms-txt-generator / Disallow: / à votre robots.txt. Détails dans les mentions légales.