Robots.txt
Standard used to advise web crawlers and scrapers not to index a web page or site
The Robots Exclusion Protocol (often referred to by the filename used to implement it, robots.txt) is a standard used by websites to indicate to visiting web crawlers and other web robots which portions of the website they are allowed to visit. The standard, developed in 1994, relies on voluntary compliance.
Nº Q80776 ★★★
Rare · Literature
Robots.txt
Standard used to advise web crawlers and scrapers not to index a web page or site
The Robots Exclusion Protocol (often referred to by the filename used to implement it, robots.txt) is a standard used by websites to indicate to visiting web crawlers and other web robots which portions of the website they are allowed to visit. The standard, developed in 1994, relies on voluntary compliance.
Last price
—
Floor price
—
7-day median
—
30-day sales
0
30-day range
—
In circulation
0
Price history
median
low – high
sales
No sales in this period
Show table
| Date | median | Low | High | sales |
|---|
Sales history
- Last sale
- —
- 30-day average
- —
- 30-day low
- —
- 30-day high
- —
- Sales 7d
- 0
- Sales 30d
- 0
No sales yet.
Anonymous sales: no buyer or seller shown. Figures count player-to-player sales only.
From Wikipedia
The Robots Exclusion Protocol (often referred to by the filename used to implement it, robots.txt) is a standard used by websites to indicate to visiting web crawlers and other web robots which portions of the website they are allowed to visit. The standard, developed in 1994, relies on voluntary compliance. Malicious bots can use the file as a directory of which pages to visit, though standards bodies discourage countering this with security through obscurity. Some archival sites ignore robots.txt. The standard was used in the 1990s to mitigate server overload. In the 2020s, websites began denying bots that collect information for generative AI. The "robots.txt" file can be used in conjunction with sitemaps, another robot inclusion standard for websites.
Text: Wikipédia, CC BY-SA 4.0. · Image: Belbury (Public domain) ·
Related cards
Web crawler
Internet bot that systematically browses the World Wide Web, typically for the purpose of Web indexing (web spidering)
Nº Q45842 ★★★
Robot Operating System
Collection of software frameworks for robot software development
Nº Q2160077 ★★★
Botnet
Collection of compromised internet-connected devices controlled by a third party
Nº Q317671 ★★
Internet bot
Kind of software application that runs automated tasks over the Internet
Nº Q191865 ★★★★
Neighbor Discovery Protocol
Protocol in the Internet Protocol Suite used with IPv6
Nº Q1547947 ★★
Security.txt
File format for posting security contact information
Nº Q65057155 ★★★★