← All cheat sheets

METAGOOFIL

Authorized use only. Offensive reference for systems you own or are explicitly permitted to test. You are responsible for staying within the law.

OVERVIEW#

Metagoofil is an information gathering tool that extracts metadata
from public documents (pdf, doc, xls, ppt) belonging to a target
domain. It searches Google for files, downloads them, and extracts
usernames, software versions, paths, and email addresses.

BASIC USAGE#

metagoofil -d <domain> -t <filetypes> -o <output_dir>
                                 # Basic metadata extraction

OPTIONS#

metagoofil -d <domain>           # Target domain
metagoofil -t <types>            # File types (comma-separated)
metagoofil -l <limit>            # Max results to search (default: 200)
metagoofil -n <download>         # Max files to download
metagoofil -o <dir>              # Output directory for downloads
metagoofil -f <file>             # Output report filename
metagoofil -w                    # Delay between requests (seconds)

FILE TYPES#

metagoofil -d example.com -t pdf      # PDF documents
metagoofil -d example.com -t doc      # Word documents
metagoofil -d example.com -t xls      # Excel spreadsheets
metagoofil -d example.com -t ppt      # PowerPoint presentations
metagoofil -d example.com -t docx     # Word (OOXML)
metagoofil -d example.com -t xlsx     # Excel (OOXML)
metagoofil -d example.com -t pptx     # PowerPoint (OOXML)
metagoofil -d example.com -t pdf,doc,xls,ppt
                                      # Multiple types

EXAMPLES#

# Extract metadata from PDFs and Word docs
metagoofil -d example.com -t pdf,doc -l 100 -n 50 -o /tmp/meta -f report.html

# Quick scan for all common file types
metagoofil -d example.com -t pdf,doc,xls,ppt,docx,xlsx,pptx -l 50 -n 25 -o /tmp/results

# Deep search with more results
metagoofil -d example.com -t pdf -l 500 -n 200 -o /tmp/deep_meta

# Targeted Excel file search
metagoofil -d example.com -t xls,xlsx -l 100 -n 50 -o /tmp/excel_meta

EXTRACTED METADATA#

# Metagoofil typically extracts:
# - Author names / usernames
# - Software and version (e.g., Microsoft Office 2016)
# - Operating system information
# - File paths (revealing internal directory structure)
# - Email addresses
# - Creation and modification dates
# - Printer names
# - Company/organization names

ANALYSIS TIPS#

# After extraction, look for:
# 1. Usernames → potential login credentials
# 2. Software versions → vulnerability research
# 3. Internal paths → internal network structure
# 4. Email formats → pattern for further enumeration
# 5. Old documents → outdated/vulnerable software

# Manual metadata extraction:
exiftool downloaded_file.pdf     # Detailed metadata with exiftool
strings downloaded_file.doc      # Raw strings from binary docs

NOTES#

- Uses Google search (may be rate-limited/blocked)
- Download directory must exist before running
- Large scans may take significant time
- Some documents may be password-protected
- Pair with theHarvester for email enumeration
- Review downloaded files for sensitive information
- Python-based tool
- Respect target scope during engagements