FOCA
Authorized use only. Offensive reference for systems you own or are explicitly permitted to test. You are responsible for staying within the law.
Metadata extraction and analysis tool. Discovers hidden information in documents (PDF, DOCX, XLSX, PPTX, etc.) found on target websites.
OVERVIEW#
FOCA (Fingerprinting Organizations with Collected Archives) extracts metadata from public documents to reveal: - Usernames and email addresses - Software versions (OS, Office, PDF creators) - Internal network paths and server names - Printer names and shared resources - Hidden revision history and comments - GPS coordinates (from images)
INSTALLATION#
# Windows only (GUI application, .NET) # Download from: https://github.com/ElevenPaths/FOCA/releases # Alternative: run on Linux via Wine or VM # Dependencies - .NET Framework 4.7.2+ - SQL Server Express (LocalDB, auto-installed)
BASIC WORKFLOW#
1. Create new project (File > New Project) 2. Enter target domain (e.g., example.com) 3. Configure search engines (Google, Bing, DuckDuckGo) 4. Select file types to search for 5. Click "Search All" to find documents 6. Download discovered documents 7. Extract metadata from downloaded files 8. Analyze results in the metadata panel
SUPPORTED FILE TYPES#
Type Extensions Metadata Extracted
---- ---------- ------------------
Office .doc, .docx Author, company, template, revision
.xls, .xlsx Printer, paths, last saved by
.ppt, .pptx Software version, comments
PDF .pdf Creator, producer, author, dates
OpenOffice .odt, .ods, .odp Author, generator, statistics
Images .jpg, .png, .gif EXIF (GPS, camera, software)
SVG .svg Editor, embedded paths
InDesign .indd Author, fonts, links
SEARCH ENGINES#
# FOCA uses search engines to discover documents # Supported: Google, Bing, DuckDuckGo, Exalead # Search dorks used internally: site:example.com filetype:pdf site:example.com filetype:docx site:example.com filetype:xlsx site:example.com filetype:pptx
MANUAL GOOGLE DORKS FOR DOCUMENT DISCOVERY#
# Use these to manually find documents if FOCA search is slow site:target.com filetype:pdf site:target.com filetype:doc OR filetype:docx site:target.com filetype:xls OR filetype:xlsx site:target.com filetype:ppt OR filetype:pptx site:target.com filetype:odt OR filetype:ods site:target.com filetype:rtf site:target.com filetype:xml site:target.com filetype:svg # With keywords site:target.com filetype:pdf "confidential" site:target.com filetype:xlsx "internal" site:target.com filetype:docx "password" site:target.com filetype:pdf "draft"
METADATA EXTRACTED#
# User Information - Author / Creator name - Last Modified By - Company name - Email addresses (sometimes embedded) # Software Information - Operating system (from paths) - Office version (e.g., Microsoft Office 16.0) - PDF creator (Acrobat, LibreOffice, LaTeX, etc.) - Application version numbers # Network Information - Internal file paths (\\SERVER\Share\folder\) - Printer names (\\PRINTER01\HP LaserJet) - UNC paths revealing server names - Mapped drive letters with server references - Local file paths (C:\Users\jsmith\Documents\) # Document Information - Creation date - Last modified date - Revision count - Print date - Template name used
FOCA ANALYSIS TABS#
# Network tab - Discovered servers and hostnames from metadata - Internal IP ranges inferred from paths - DNS resolution of discovered hostnames # Metadata tab - Per-document breakdown of all extracted metadata - Users, dates, software, paths, printers # Users tab - Aggregated list of all discovered usernames - Cross-referenced across documents - Useful for password spraying username lists # Software tab - Software versions found across documents - Identifies outdated/vulnerable versions # Servers tab - Internal servers discovered from UNC paths - Printer servers
COMMAND-LINE ALTERNATIVES#
# exiftool (cross-platform, most comprehensive) exiftool document.pdf # All metadata exiftool -a -u -g1 document.pdf # All tags, grouped exiftool -Author -Creator -Producer doc.pdf # Specific fields exiftool -r -ext pdf /path/to/docs/ # Recursive directory exiftool -csv *.pdf > metadata.csv # CSV export exiftool -json *.docx > metadata.json # JSON export # Extract all metadata from multiple files exiftool -r -ext pdf -ext docx -ext xlsx . | grep -i "author\|creator\|company\|software" # mat2 (metadata removal tool, also useful for viewing) mat2 --show document.pdf # Show metadata mat2 document.pdf # Strip metadata # pdfinfo (PDF-specific) pdfinfo document.pdf # olemeta (Office documents, part of oletools) pip install oletools olemeta document.doc oleid document.doc # Identify features # metagoofil (automated document finder + extractor) metagoofil -d example.com -t pdf,doc,xls -o /tmp/output -l 100
EXIFTOOL QUICK REFERENCE#
# Most common metadata fields exiftool -Author file # Document author exiftool -Creator file # Creator application exiftool -Producer file # PDF producer exiftool -Company file # Company name exiftool -LastModifiedBy file # Last editor exiftool -CreateDate file # Creation date exiftool -ModifyDate file # Last modified exiftool -Software file # Software used exiftool -Title file # Document title exiftool -Subject file # Subject/description exiftool -Keywords file # Keywords exiftool -GPSPosition file # GPS coordinates (images) # Batch operations exiftool -Author -Creator *.pdf # All PDFs in directory exiftool -r -Author -csv *.docx > users.csv # Recursive CSV export
USING METADATA FOR PENTESTING#
# 1. Username enumeration # Extract authors → build username list for password spraying # Common formats: jsmith, john.smith, j.smith, smithj # 2. Software version identification # Old Office/Acrobat versions → client-side exploits # PDF creators reveal server-side tools # 3. Internal network mapping # UNC paths → server names, share names # File paths → drive mappings, directory structure # 4. Email address discovery # Some docs embed email in author/properties # Cross-reference with other OSINT sources # 5. Template analysis # Template paths reveal internal infrastructure # Custom templates may be downloadable
AUTOMATING METADATA COLLECTION#
#!/bin/bash
# Download all PDFs from a target and extract metadata
mkdir -p /tmp/foca_output
cd /tmp/foca_output
# Download documents (using wget or curl)
wget -r -l1 -A "*.pdf,*.docx,*.xlsx,*.pptx" \
--no-parent https://target.com/ -P downloads/
# Extract all metadata
exiftool -r -csv downloads/ > all_metadata.csv
# Extract just usernames
exiftool -r -Author -LastModifiedBy downloads/ | \
grep -v "^=" | grep -v "^-" | sort -u > users.txt
TIPS#
- Always check multiple file types (PDF, Office, images) - UNC paths are gold for internal network mapping - Old documents often have more metadata (less scrubbing) - Cross-reference discovered usernames with LinkedIn - Government and education sites often have rich metadata - exiftool is the best CLI alternative to FOCA - metagoofil automates the download + extraction workflow - mat2 can strip metadata before sharing your own documents - Check revision history in Office docs for deleted content - GPS data in images can reveal physical locations