METAGOOFIL
Authorized use only. Offensive reference for systems you own or are explicitly permitted to test. You are responsible for staying within the law.
OVERVIEW#
Metagoofil is an information gathering tool that extracts metadata from public documents (pdf, doc, xls, ppt) belonging to a target domain. It searches Google for files, downloads them, and extracts usernames, software versions, paths, and email addresses.
BASIC USAGE#
metagoofil -d <domain> -t <filetypes> -o <output_dir>
# Basic metadata extraction
OPTIONS#
metagoofil -d <domain> # Target domain metagoofil -t <types> # File types (comma-separated) metagoofil -l <limit> # Max results to search (default: 200) metagoofil -n <download> # Max files to download metagoofil -o <dir> # Output directory for downloads metagoofil -f <file> # Output report filename metagoofil -w # Delay between requests (seconds)
FILE TYPES#
metagoofil -d example.com -t pdf # PDF documents
metagoofil -d example.com -t doc # Word documents
metagoofil -d example.com -t xls # Excel spreadsheets
metagoofil -d example.com -t ppt # PowerPoint presentations
metagoofil -d example.com -t docx # Word (OOXML)
metagoofil -d example.com -t xlsx # Excel (OOXML)
metagoofil -d example.com -t pptx # PowerPoint (OOXML)
metagoofil -d example.com -t pdf,doc,xls,ppt
# Multiple types
EXAMPLES#
# Extract metadata from PDFs and Word docs metagoofil -d example.com -t pdf,doc -l 100 -n 50 -o /tmp/meta -f report.html # Quick scan for all common file types metagoofil -d example.com -t pdf,doc,xls,ppt,docx,xlsx,pptx -l 50 -n 25 -o /tmp/results # Deep search with more results metagoofil -d example.com -t pdf -l 500 -n 200 -o /tmp/deep_meta # Targeted Excel file search metagoofil -d example.com -t xls,xlsx -l 100 -n 50 -o /tmp/excel_meta
EXTRACTED METADATA#
# Metagoofil typically extracts: # - Author names / usernames # - Software and version (e.g., Microsoft Office 2016) # - Operating system information # - File paths (revealing internal directory structure) # - Email addresses # - Creation and modification dates # - Printer names # - Company/organization names
ANALYSIS TIPS#
# After extraction, look for: # 1. Usernames → potential login credentials # 2. Software versions → vulnerability research # 3. Internal paths → internal network structure # 4. Email formats → pattern for further enumeration # 5. Old documents → outdated/vulnerable software # Manual metadata extraction: exiftool downloaded_file.pdf # Detailed metadata with exiftool strings downloaded_file.doc # Raw strings from binary docs
NOTES#
- Uses Google search (may be rate-limited/blocked) - Download directory must exist before running - Large scans may take significant time - Some documents may be password-protected - Pair with theHarvester for email enumeration - Review downloaded files for sensitive information - Python-based tool - Respect target scope during engagements