← All cheat sheets

PARSERO

Authorized use only. Offensive reference for systems you own or are explicitly permitted to test. You are responsible for staying within the law.

OVERVIEW#

Parsero reads the robots.txt file of a web server and looks at the
Disallow entries. It then tests those paths with HTTP requests to
check if they are accessible (HTTP 200) despite being "disallowed",
revealing potentially sensitive content.

BASIC USAGE#

parsero -u <url>                 # Check robots.txt disallowed paths
parsero -u https://example.com   # HTTPS target

OPTIONS#

parsero -u <url>                 # Target URL
parsero -o <file>                # Save results to file
parsero -sb                      # Only show HTTP 200 results

EXAMPLES#

# Check all disallowed paths
parsero -u https://example.com

# Only show accessible (200 OK) paths
parsero -u https://example.com -sb

# Save results to file
parsero -u https://example.com -o results.txt

# Check with only accessible results saved
parsero -u https://example.com -sb -o accessible.txt

INTERPRETING OUTPUT#

# Output shows each Disallow path with HTTP status:
#
# [200] https://example.com/admin/
#   → Accessible! Potentially sensitive page exposed
#
# [403] https://example.com/private/
#   → Forbidden (properly restricted)
#
# [404] https://example.com/old-page/
#   → Not found (path no longer exists)
#
# [301/302] https://example.com/redirect/
#   → Redirects to another location
#
# [500] https://example.com/error/
#   → Server error (may be interesting)

WHAT TO LOOK FOR#

# Accessible paths (HTTP 200) that are disallowed:
# - /admin/          - Admin panels
# - /backup/         - Backup files
# - /config/         - Configuration files
# - /database/       - Database files/exports
# - /logs/           - Log files
# - /tmp/            - Temporary files
# - /api/            - API endpoints
# - /test/           - Test environments
# - /.git/           - Git repository data
# - /.svn/           - SVN repository data
# - /wp-admin/       - WordPress admin
# - /phpmyadmin/     - Database management

MANUAL ALTERNATIVE#

# View robots.txt manually
curl https://example.com/robots.txt

# Check individual paths
curl -s -o /dev/null -w "%{http_code}" https://example.com/admin/

NOTES#

- robots.txt is publicly accessible by design
- Disallow ≠ access control (only a suggestion to crawlers)
- HTTP 200 on disallowed paths = information disclosure
- Many sites accidentally expose sensitive directories
- Python-based tool
- Useful as a first step in web application recon
- Pair with directory brute forcing (gobuster, dirbuster)