Results 1 to 8 of 8
  1. #1
    Member
    Join Date
    January 18th, 2005
    Posts
    189
    Hello -

    Just wanted to check and make sure this robot.txt file will work, and I didn't do something wrong that would prevent my site from being crawled...

    Here is the contents of my robot.txt file in the parent (root) directory of my website.

    User-agent: *
    Disallow: /dynamic/trackster.htm

    Now, this robot.txt file will prevent SE's which abide by it from crawling to any page on my site with /dynamic/trackster.htm in the URL. Correct? It is my link tracking page, which tracks clicks, and it needn't be spidered...

    Dan

  2. #2
    ABW Ambassador
    Join Date
    January 18th, 2005
    Location
    United Kingdom
    Posts
    1,797
    Well, the format is right but its a robots.txt file not robot.txt

    Search Engine Positioning - 1 Design 4 Life

  3. #3
    Member
    Join Date
    January 18th, 2005
    Posts
    189
    Heh...oops, yeah I knew that, didn't type it that way in the above though..

    Thanks,

    Dan

  4. #4
    Member
    Join Date
    January 18th, 2005
    Posts
    110
    1) That line will only exclude the single page domainroot/dynamic/trackster.htm, as the robots.txt file MUST reside in the domain root, and all URLs are considered relative to it. Is that a problem? Do you have multiple /dynamic/trackster.htm files?

    2) The lines should have UNIX line endings apparently, to conform exactly to the spec. Some dodgy old robots do need this, or they misinterpret the file. A decent text editor should allow you to specify this on save (it may well be true anyway)

  5. #5
    Member
    Join Date
    January 18th, 2005
    Posts
    189
    Troll -

    Thanks for the input.

    1) In reality, the URL to the link will be "http://mydomain.com/dynamic/trackster.htm?page=1" - the robots.txt above will tell google and the others to disregard any link using the root domain with "/dynamic/trackster.htm" correct?

    2) UNIX line endings? Interesting...I'll throw them in there...

    Thanks,

    Dan

  6. #6
    ABW Ambassador
    Join Date
    January 18th, 2005
    Location
    United Kingdom
    Posts
    1,797
    <BLOCKQUOTE class="ip-ubbcode-quote"><font size="-1">quote:</font><HR> 1) In reality, the URL to the link will be "http://mydomain.com/dynamic/trackster.htm?page=1" - the robots.txt above will tell google and the others to disregard any link using the root domain with "/dynamic/trackster.htm" correct?
    <HR></BLOCKQUOTE>
    Ah, OK. If you have pages like /dynamic/trackster.html?page=4 etc, then you will need the following to disallow all the trackster pages:

    User-agent: *
    Disallow: /dynamic/trackster

    That will disallow everything in the dynamic folder that starts with trackster.

    Search Engine Positioning - 1 Design 4 Life

  7. #7
    Member
    Join Date
    January 18th, 2005
    Posts
    189
    Mark -

    Thanks a lot.

    Dan

  8. #8
    2005 Linkshare Golden Link Award Winner  ecomcity's Avatar
    Join Date
    January 18th, 2005
    Location
    St Clair Shores MI.
    Posts
    17,328
    Does this robots.txt file appear properly formated?

    User-agent:*
    Disallow:/stats/
    Disallow:/_private/
    Disallow:/_borders/
    Disallow:/_fpclass/
    Disallow:/_overlay/
    Disallow:/_themes/
    Disallow:/_vti_bin/
    Disallow:/_vti_cnf/
    Disallow:/_vti_log/
    Disallow:/_vti_pvt/
    Disallow:/_vti_txt/
    Disallow:/images/
    Disallow:/club/
    User-agent: TurnitinBot
    Disallow: /
    User-agent: scooter
    Disallow: /

    User-agent: grub-client
    Disallow: /

    User-agent: grub
    Disallow: /

    User-agent: looksmart
    Disallow: /

    User-agent: WebZip
    Disallow: /

    User-agent: larbin
    Disallow: /

    User-agent: b2w/0.1
    Disallow: /

    User-agent: Copernic
    Disallow: /

    User-agent: psbot
    Disallow: /

    User-agent: Python-urllib
    Disallow: /

    User-agent: Googlebot-Image
    Disallow: /

    User-agent: NetMechanic
    Disallow: /

    User-agent: URL_Spider_Pro
    Disallow: /

    User-agent: CherryPicker
    Disallow: /

    User-agent: EmailCollector
    Disallow: /

    User-agent: EmailSiphon
    Disallow: /

    User-agent: WebBandit
    Disallow: /

    User-agent: EmailWolf
    Disallow: /

    User-agent: ExtractorPro
    Disallow: /

    User-agent: CopyRightCheck
    Disallow: /

    User-agent: Crescent
    Disallow: /

    User-agent: SiteSnagger
    Disallow: /

    User-agent: ProWebWalker
    Disallow: /

    User-agent: CheeseBot
    Disallow: /

    User-agent: LNSpiderguy
    Disallow: /

    User-agent: Mozilla
    Disallow: /

    User-agent: mozilla
    Disallow: /

    User-agent: mozilla/3
    Disallow: /

    User-agent: mozilla/4
    Disallow: /

    User-agent: mozilla/5
    Disallow: /

    User-agent: Mozilla/4.0 (compatible; MSIE 4.0; Windows NT)
    Disallow: /

    User-agent: Mozilla/4.0 (compatible; MSIE 4.0; Windows 95)
    Disallow: /

    User-agent: Mozilla/4.0 (compatible; MSIE 4.0; Windows 98)
    Disallow: /

    User-agent: Mozilla/4.0 (compatible; MSIE 4.0; Windows XP)
    Disallow: /

    User-agent: Mozilla/4.0 (compatible; MSIE 4.0; Windows 2000)
    Disallow: /

    User-agent: ia_archiver
    Disallow: /

    User-agent: ia_archiver/1.6
    Disallow: /

    User-agent: Alexibot
    Disallow: /

    User-agent: Teleport
    Disallow: /

    User-agent: TeleportPro
    Disallow: /

    User-agent: MIIxpc
    Disallow: /

    User-agent: Telesoft
    Disallow: /

    User-agent: Website Quester
    Disallow: /

    User-agent: moget/2.1
    Disallow: /

    User-agent: WebZip/4.0
    Disallow: /

    User-agent: WebStripper
    Disallow: /

    User-agent: WebSauger
    Disallow: /

    User-agent: WebCopier
    Disallow: /

    User-agent: NetAnts
    Disallow: /

    User-agent: Mister PiX
    Disallow: /

    User-agent: WebAuto
    Disallow: /

    User-agent: TheNomad
    Disallow: /

    User-agent: WWW-Collector-E
    Disallow: /

    User-agent: RMA
    Disallow: /

    User-agent: libWeb/clsHTTP
    Disallow: /

    User-agent: asterias
    Disallow: /

    User-agent: httplib
    Disallow: /

    User-agent: turingos
    Disallow: /

    User-agent: spanner
    Disallow: /

    User-agent: InfoNaviRobot
    Disallow: /

    User-agent: Harvest/1.5
    Disallow: /

    User-agent: Bullseye/1.0
    Disallow: /

    User-agent: Mozilla/4.0 (compatible; BullsEye; Windows 95)
    Disallow: /

    User-agent: Crescent Internet ToolPak HTTP OLE Control v.1.0
    Disallow: /

    User-agent: CherryPickerSE/1.0
    Disallow: /

    User-agent: CherryPickerElite/1.0
    Disallow: /

    User-agent: WebBandit/3.50
    Disallow: /

    User-agent: NICErsPRO
    Disallow: /

    User-agent: Microsoft URL Control - 5.01.4511
    Disallow: /

    User-agent: DittoSpyder
    Disallow: /

    User-agent: Foobot
    Disallow: /

    User-agent: WebmasterWorldForumBot
    Disallow: /

    User-agent: SpankBot
    Disallow: /

    User-agent: BotALot
    Disallow: /

    User-agent: lwp-trivial/1.34
    Disallow: /

    User-agent: lwp-trivial
    Disallow: /

    User-agent: BunnySlippers
    Disallow: /

    User-agent: Microsoft URL Control - 6.00.8169
    Disallow: /

    User-agent: URLy Warning
    Disallow: /

    User-agent: Wget/1.6
    Disallow: /

    User-agent: Wget/1.5.3
    Disallow: /

    User-agent: Wget
    Disallow: /

    User-agent: LinkWalker
    Disallow: /

    User-agent: cosmos
    Disallow: /

    User-agent: moget
    Disallow: /

    User-agent: hloader
    Disallow: /

    User-agent: humanlinks
    Disallow: /

    User-agent: LinkextractorPro
    Disallow: /

    User-agent: Offline Explorer
    Disallow: /

    User-agent: Mata Hari
    Disallow: /

    User-agent: LexiBot
    Disallow: /

    User-agent: Web Image Collector
    Disallow: /

    User-agent: The Intraformant
    Disallow: /

    User-agent: True_Robot/1.0
    Disallow: /

    User-agent: True_Robot
    Disallow: /

    User-agent: BlowFish/1.0
    Disallow: /

    User-agent: JennyBot
    Disallow: /

    User-agent: MIIxpc/4.2
    Disallow: /

    User-agent: BuiltBotTough
    Disallow: /

    User-agent: ProPowerBot/2.14
    Disallow: /

    User-agent: BackDoorBot/1.0
    Disallow: /

    User-agent: toCrawl/UrlDispatcher
    Disallow: /

    User-agent: WebEnhancer
    Disallow: /

    User-agent: suzuran
    Disallow: /

    User-agent: VCI WebViewer VCI WebViewer Win32
    Disallow: /

    User-agent: VCI
    Disallow: /

    User-agent: Szukacz/1.4
    Disallow: /

    User-agent: QueryN Metasearch
    Disallow: /

    User-agent: Openfind data gathere
    Disallow: /

    User-agent: Openfind
    Disallow: /

    User-agent: Xenu's Link Sleuth 1.1c
    Disallow: /

    User-agent: Xenu's
    Disallow: /

    User-agent: Zeus
    Disallow: /

    User-agent: RepoMonkey Bait & Tackle/v1.01
    Disallow: /

    User-agent: RepoMonkey
    Disallow: /

    User-agent: Microsoft URL Control
    Disallow: /

    User-agent: Openbot
    Disallow: /

    User-agent: URL Control
    Disallow: /

    User-agent: Zeus Link Scout
    Disallow: /

    User-agent: Zeus 32297 Webster Pro V2.9 Win32
    Disallow: /

    User-agent: Webster Pro
    Disallow: /

    User-agent: EroCrawler
    Disallow: /

    User-agent: LinkScan/8.1a Unix
    Disallow: /

    User-agent: Keyword Density/0.9
    Disallow: /

    User-agent: Kenjin Spider
    Disallow: /

    User-agent: Iron33/1.0.2
    Disallow: /

    User-agent: Bookmark search tool
    Disallow: /

    User-agent: GetRight/4.2
    Disallow: /

    User-agent: FairAd Client
    Disallow: /

    User-agent: Gaisbot
    Disallow: /

    User-agent: Aqua_Products
    Disallow: /

    User-agent: Radiation Retriever 1.1
    Disallow: /

    User-agent: WebmasterWorld Extractor
    Disallow: /

    User-agent: Flaming AttackBot
    Disallow: /

    User-agent: Oracle Ultra Search
    Disallow: /

    User-agent: PerMan
    Disallow: /

    User-agent: searchpreview
    Disallow: /

    Mike & Charlie ...

    If they won't adopt and feed a bird ..flip them one! BBQ some Gator and remember to flush WhenU..

  9. Newsletter Signup

+ Reply to Thread

Similar Threads

  1. Google Robot Searching for robot.txt
    By ahmar in forum Search Engine Optimization
    Replies: 4
    Last Post: December 26th, 2004, 01:26 PM
  2. Robot.txt versus Amazon.PL
    By beggers in forum Cusimano.com Scripts
    Replies: 12
    Last Post: March 18th, 2003, 04:16 PM
  3. A Lesson About Robot.txt
    By seaslug44 in forum Search Engine Optimization
    Replies: 1
    Last Post: February 22nd, 2002, 07:49 PM
  4. Robot txt
    By mousejockey in forum Programming / Datafeeds / Tools
    Replies: 12
    Last Post: January 14th, 2002, 05:05 PM

Posting Permissions

  • You may not post new threads
  • You may not post replies
  • You may not post attachments
  • You may not edit your posts
  •