Skip to content Hivex tools · Website Free · No account
Hivex / Tools / robots.txt tester / http://nytimes.com/ robots.txt of nytimes.com. See which search engines and AI crawlers may fetch a page, and the exact line of robots.txt that decides it.
How the rules are read. A crawler looks for groups with its own name in User-agent:, case does not matter, and combines them; with none, it uses the * group. Of the Allow: and Disallow: lines that match the path, the longest wins. A tie goes to Allow. * matches any run of characters and $ the end of the address. No file (404) means everything may be fetched; a server error (5xx) means crawlers treat everything as blocked. /robots.txt itself is always allowed. AI crawlers, by what they do. Training : GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Gemini), CCBot (Common Crawl), Applebot-Extended. Blocking them asks that your pages not be used to train models.Search : OAI-SearchBot, Claude-SearchBot, PerplexityBot. Blocking them can keep your site out of those assistants' answers.On a user's request : ChatGPT-User and Claude-User fetch a page someone asked about; OpenAI notes robots.txt may not apply to those.Blocked is not hidden. robots.txt stops crawling, not indexing: a blocked page can still show in search results, without a description, if other pages link to it. To keep a page out of results, let it be crawled and add noindex. Never rely on robots.txt for anything private; the file itself is public.
Questions people ask. How does robots.txt decide whether a page is blocked? A crawler uses the group for its own name, or the * group if it has none. Of the Allow and Disallow lines that match the path, the longest one wins; when an Allow and a Disallow are equally long, Allow wins (RFC 9309).
How do I block AI crawlers but keep Google? Give each AI crawler its own group with Disallow: /, for example User-agent: GPTBot, User-agent: ClaudeBot, User-agent: Google-Extended, User-agent: CCBot. Googlebot keeps following its own rules, so search is unaffected.
Does Disallow keep a page out of Google? No. It stops crawling, but a blocked page can still appear in results if other sites link to it. To keep a page out, allow crawling and add a noindex meta tag or header.
What happens if robots.txt is missing or broken? A missing file (404) means crawlers may fetch everything. A server error (5xx) makes crawlers treat the whole site as blocked until it is fixed.
Is there a size limit? Crawlers must read at least 500 KiB, and Google stops there: rules further down are ignored.
Hivex index
Short names, still free to register.
Starting something new? Hivex keeps a live index of short, brandable .si names nobody has claimed yet, each checked with the registry.
Browse free names From code, or an AI assistant. More Website tools Redirect checker Every hop of a URL: status code, target and time. HTTP header checker All response headers, with the security headers checked. Security headers checker A grade for HSTS, CSP, framing and cookies, with fixes. SSL checker Trust, expiry, the chain and TLS versions of any site. Open Graph checker How a link looks when shared, with the tags checked. Is it down? Down for everyone or just you, and where it fails. Sitemap checker Limits, dates and hosts checked, a sample requested live. Who hosts this website Web host, CDN, DNS host, mail provider and registrar. Tech stack checker CMS, framework, hosting and trackers of any site. Domain checker Is a name free to register? Checked live on any extension. All free tools Partners
Domains worth owning.
Checked live · No account needed · Premium sales with AtomPay protection
© 2026 hivex.si. Availability changes all the time, so confirm at the registrar before you pay. Not affiliated with ARNES or register.si. Some domain and product links earn Hivex a commission, at no extra cost to you.
http://nytimes.com/robots.txt
Checked 03:17:10 UTC · Fetched from Hivex's servers
Tested path /
Groups 63
Sitemaps 25
Size 8.3 KiB Note: Every AI crawler listed below is blocked from /.Good: Google may crawl /.Who may fetch / RFC 9309 rules: longest match wins
Crawler / Decided by Search engines Googlebot Google Search Allowed no rule matches (group Googlebot) Bingbot Bing (and search built on it) Allowed no rule matches (group *) AI training Google-Extended Google: Gemini training and grounding Blocked line 230: Disallow: / GPTBot OpenAI: model training Blocked line 233: Disallow: /
To test another page, put its path in the address you check, such as example.com/private/page.
The file Redirected to https://www.nytimes.com/robots.txt
1 # New York Times content is made available for your personal, non-commercial 2 # use subject to our Terms of Service here: 3 # https://help.nytimes.com/hc/en-us/articles/115014893428-Terms-of-Service. 4 # Use of any device, tool, or process designed to data mine or scrape the content 5 # using automated means is prohibited without prior written permission from 6 # The New York Times Company. Prohibited uses include but are not limited to: 7 # (1) text and data mining activities under Art. 4 of the EU Directive on Copyright in 8 # the Digital Single Market; 9 # (2) the development of any software, machine learning, artificial intelligence (AI), ClaudeBot Anthropic: model training Blocked line 180: Disallow: /
CCBot Common Crawl: open dataset many models train on Blocked line 174: Disallow: /
Applebot-Extended Apple: AI training Blocked line 152: Disallow: /
AI search OAI-SearchBot OpenAI: ChatGPT search Blocked line 266: Disallow: / Claude-SearchBot Anthropic: Claude's search Blocked line 183: Disallow: / PerplexityBot Perplexity: search index Blocked line 279: Disallow: /
AI, fetched for a user ChatGPT-User OpenAI: pages a ChatGPT user asks for Blocked line 177: Disallow: / Claude-User Anthropic: pages a Claude user asks for Blocked line 186: Disallow: /
10 # and/or large language models (LLMs);
11 # (3) creating or providing archived or cached data sets containing our content to others; and/or
12 # (4) any commercial purposes.
13 # Contact https://nytlicensing.com/contact/ for assistance.
14
15 User-agent: *
16 User-agent: Googlebot
17 Disallow: /ads/
18 Disallow: /adx/bin/
19 Disallow: /athletic/wp/wp-admin/
20 Allow: /athletic/wp/wp-admin/admin-ajax.php
21 Disallow: /athletic/async-*
22 Disallow: /athletic/search/*
23 Allow: /athletic/search/$
24 Disallow: /athletic/checkout/
25 Disallow: /athletic/checkout?plan_id*
26 Allow: /athletic/checkout/$
27 Disallow: /athletic/checkout2*
28 Disallow: /athletic/login/
29 Disallow: /athletic/login?login_source*
30 Disallow: /athletic/login?ref_page*
31 Allow: /athletic/login/$
32 Disallow: /athletic/login2/
33 Disallow: /athletic/login2?login_source*
34 Disallow: /athletic/login2?ref_page*
35 Allow: /athletic/login2/$
36 Disallow: /athletic/report/
37 Disallow: /athletic/*/discuss/*
38 Allow: /athletic/live-blogs/discuss/
39 Allow: /athletic/mlb/game/discuss/
40 Disallow: /athletic/register/
41 Disallow: /athletic/register?welcome_redirect*
42 Disallow: /athletic/register2/
43 Disallow: /athletic/register2?welcome_redirect*
44 Disallow: /athletic/betmgm-redirect*
45 Disallow: /athletic/cdn-cgi/
46 Disallow: /athletic/verizon/*
47 Disallow: /athletic/forgot-password/*
48 Disallow: /athletic/forgot-password2/*
49 Disallow: /athletic/amp-social-login*
50 Disallow: /athletic/track-analytics/
51 Disallow: /athletic/amp-auth/
52 Disallow: /athletic/rss-feed/
53 Disallow: /athletic/*?*rss=1
54 Disallow: /athletic/global-color-test.php
55 Disallow: /athletic/global-font-test.php
56 Disallow: /athletic/graphql*
57 Disallow: /athletic/api*
58 Disallow: /athletic/ip*
59 Disallow: /athletic/call-set-cookie-with-context/*
60 Disallow: /athletic/get-current-user/
61 Disallow: /athletic/pv.json
62 Disallow: /athletic/following-feed-test/*
63 Disallow: /athletic*/boxscore/*
64 Disallow: /athletic/feed-test/
65 Disallow: /athletic*/signed-mp3-redirect-url/*
66 Disallow: /athletic/embedded-interactive/*
67 Disallow: /card/panel/
68 Disallow: /panel/
69 Disallow: /puzzles/leaderboards/invite/*
70 Disallow: /svc
71 Allow: /svc/crosswords
72 Allow: /svc/games
73 Allow: /svc/letter-boxed
74 Allow: /svc/spelling-bee
75 Allow: /svc/wordle
76 Allow: /svc/connections
77 Allow: /svc/sudoku
78 Allow: /svc/strands
79 Allow: /svc/pips
80 Disallow: /video/embedded/*
81 Disallow: /search
82 Disallow: /multiproduct/
83 Disallow: /hd/
84 Disallow: /inyt/
85 Disallow: /*?*query=
86 Disallow: /*.pdf$
87 Disallow: /*?*login=
88 Disallow: /*?*campaignId=
89 Disallow: /*?*mcubz=
90 Disallow: /*?*smprod=
91 Disallow: /*?*ProfileID=
92 Disallow: /*?*ListingID=
93 Disallow: /*?*campaign_id=
94 Disallow: /*?*hybrid=
95 Disallow: /*?*entry=
96 Disallow: /*?*embed=
97 Disallow: /*?ls=
98 Disallow: /*?*&ls=
99 Disallow: /v1/initialize*
100 Disallow: /v1/rgstr*
101 Disallow: /wirecutter/wp-admin/
102 Disallow: /wirecutter/*.zip$
103 Disallow: /wirecutter/*.csv$
104 Disallow: /wirecutter/deals/beta
105 Disallow: /wirecutter/data-requests
106 Disallow: /wirecutter/search
107 Disallow: /wirecutter/*?s=
108 Disallow: /wirecutter/*&xid=
109 Disallow: /wirecutter/*?q=
110 Disallow: /wirecutter/*?l=
111 Disallow: /wirecutter/*?merchant=
112 Disallow: /wirecutter/out/
113 Disallow: /search
114 Disallow: /subscription/*?*source=
115 Disallow: /subscription/*?*onboarded=
116 Disallow: /*?*smid=
117 Disallow: /*?*partner=
118 Disallow: /*?*utm_source=
119 Allow: /wirecutter/*?*utm_source=
120 Allow: /ads/public/
121 Allow: /svc/news/v3/all/pshb.rss
122 Allow: /wirecutter/reviews/*?*utm_source=
123 Allow: /wirecutter/blog/*?*utm_source=
124 Allow: /wirecutter/lists/*?*utm_source=
125 Allow: /wirecutter/gifts/*?*utm_source=
126
127
128 # Googlebot Specific Rules
129
130 User-agent: Googlebot
131 Disallow: /athletic*adgroupid*
132 Disallow: /athletic*campaignid*
133 Disallow: /athletic*ad_id*
134 Disallow: /athletic*access_token*
135 Disallow: /athletic*amp_reader_id*
136 Disallow: /athletic*/?source=*
137 Disallow: /athletic/*?*embed=1
138
139
140 # Disallow Rules
141
142 User-agent: AliyunSecBot
143 Disallow: /
144
145 User-agent: Amazonbot
146 Disallow: /wirecutter/
147
148 User-agent: anthropic-ai
149 Disallow: /
150
151 User-agent: Applebot-Extended
152 Disallow: /
153
154 User-agent: archive.org_bot
155 Disallow: /
156
157 User-agent: AudigentAdBot
158 Disallow: /
159
160 User-agent: AwarioRssBot
161 User-agent: AwarioSmartBot
162 Disallow: /
163
164 User-agent: BLEXBot
165 Disallow: /
166
167 User-agent: Brightbot
168 Disallow: /
169
170 User-agent: Bytespider
171 Disallow: /
172
173 User-agent: CCBot
174 Disallow: /
175
176 User-agent: ChatGPT-User
177 Disallow: /
178
179 User-agent: ClaudeBot
180 Disallow: /
181
182 User-agent: Claude-SearchBot
183 Disallow: /
184
185 User-agent: Claude-User
186 Disallow: /
187
188 User-agent: Claude-Web
189 Disallow: /
190
191 User-agent: cohere-ai
192 Disallow: /
193
194 User-agent: DataForSeoBot
195 Disallow: /
196
197 User-agent: Diffbot
198 Disallow: /
199
200 User-agent: Diffbot-User
201 Disallow: /
202
203 User-agent: DuckAssistBot
204 Disallow: /
205
206 User-agent: EchoboxBot
207 Disallow: /
208
209 User-agent: ExaBot
210 Disallow: /
211
212 User-agent: ExaSearchBot
213 Disallow: /
214
215 User-agent: FacebookBot
216 Disallow: /
217
218 User-agent: FirecrawlAgent
219 Disallow: /
220
221 User-agent: FriendlyCrawler
222 Disallow: /
223
224 User-agent: Google-CloudVertexBot
225 Disallow: /
226 Allow: /wirecutter/
227 Allow: /athletic/
228
229 User-agent: Google-Extended
230 Disallow: /
231
232 User-agent: GPTBot
233 Disallow: /
234
235 User-agent: ImagesiftBot
236 Disallow: /
237
238 User-agent: Jetslide
239 Disallow: /
240
241 User-agent: magpie-crawler
242 Disallow: /
243
244 User-agent: Meta-ExternalAgent
245 User-agent: meta-externalagent
246 Disallow: /
247
248 User-agent: Meta-ExternalFetcher
249 User-agent: meta-externalfetcher
250 Disallow: /
251
252 User-agent: Meta-WebIndexer
253 User-agent: meta-webindexer/1.1
254 Disallow: /
255
256 User-agent: MyCentralAIScraperBot
257 Disallow: /
258
259 User-agent: NewsNow
260 Disallow: /
261
262 User-agent: news-please
263 Disallow: /
264
265 User-agent: OAI-SearchBot
266 Disallow: /
267
268 User-agent: omgili
269 Disallow: /
270
271 User-agent: omgilibot
272 Disallow: /
273
274 User-agent: peer39_crawler
275 User-agent: peer39_crawler/1.0
276 Disallow: /
277
278 User-agent: PerplexityBot
279 Disallow: /
280
281 User-agent: Perplexity-User
282 Disallow: /
283
284 User-agent: Poseidon Research Crawler
285 Disallow: /
286
287 User-agent: quillbot.com
288 Disallow: /
289
290 User-agent: Quora-Bot
291 Disallow: /
292
293 User-agent: Scrapy
294 Disallow: /
295
296 User-agent: SeekrBot
297 Disallow: /
298
299 User-agent: SeznamHomepageCrawler
300 Disallow: /
301
302 User-agent: ShapBot
303 Disallow: /
304
305 User-agent: TaraGroup Intelligent Bot
306 Disallow: /
307
308 User-agent: Timpibot
309 Disallow: /
310
311 User-agent: TurnitinBot
312 Disallow: /
313
314 User-agent: ViennaTinyBot
315 Disallow: /
316
317 User-agent: YouBot
318 Disallow: /
319
320
321 # Ad Bot Rules
322
323 User-agent: AmazonAdBot
324 Allow: /
325
326
327 # Social Bot Rules
328
329 User-agent: facebookexternalhit
330 Allow: /*?*smid=
331
332 User-agent: Twitterbot
333 Allow: /*?*smid=
334
335 User-agent: RedditBot
336 Allow: /*?*smid=
337
338 # Sitemaps
339
340 Sitemap: https://www.nytimes.com/sitemaps/new/news.xml.gz
341 Sitemap: https://www.nytimes.com/sitemaps/new/sitemap.xml.gz
342 Sitemap: https://www.nytimes.com/sitemaps/new/collections.xml.gz
343 Sitemap: https://www.nytimes.com/sitemaps/new/video.xml.gz
344 Sitemap: https://www.nytimes.com/sitemaps/new/cooking.xml.gz
345 Sitemap: https://www.nytimes.com/sitemaps/new/recipe-collects.xml.gz
346 Sitemap: https://www.nytimes.com/sitemaps/new/regions.xml
347 Sitemap: https://www.nytimes.com/sitemaps/new/best-sellers.xml
348 Sitemap: https://www.nytimes.com/sitemaps/new/subscription-landing-pages.xml
349 Sitemap: https://www.nytimes.com/sitemaps/new/weather.xml.gz
350 Sitemap: https://www.nytimes.com/sitemaps/new/espanol.xml.gz
351 Sitemap: https://www.nytimes.com/sitemaps/new/espanol-collects.xml.gz
352 Sitemap: https://www.nytimes.com/wirecutter/sitemapindex.xml
353 Sitemap: https://www.nytimes.com/athletic/sitemap-videos-index.xml
354 Sitemap: https://www.nytimes.com/athletic/sitemap-authors.xml
355 Sitemap: https://www.nytimes.com/athletic/sitemap-verticals.xml
356 Sitemap: https://www.nytimes.com/athletic/sitemap-teams.xml
357 Sitemap: https://www.nytimes.com/athletic/sitemap-cities.xml
358 Sitemap: https://www.nytimes.com/athletic/sitemap-players.xml
359 Sitemap: https://www.nytimes.com/athletic/sitemap-tags.xml
360 Sitemap: https://www.nytimes.com/athletic/sitemap-stats.xml
361 Sitemap: https://www.nytimes.com/athletic/sitemap-schedule.xml
362 Sitemap: https://www.nytimes.com/athletic/sitemap-roster.xml
363 Sitemap: https://www.nytimes.com/athletic/sitemap.xml
364 Sitemap: https://www.nytimes.com/games-assets/v2/assets/sitemap/games.xml
365