← All cheat sheets

FOCA

Authorized use only. Offensive reference for systems you own or are explicitly permitted to test. You are responsible for staying within the law.

Metadata extraction and analysis tool. Discovers hidden information
in documents (PDF, DOCX, XLSX, PPTX, etc.) found on target websites.

OVERVIEW#

FOCA (Fingerprinting Organizations with Collected Archives) extracts
metadata from public documents to reveal:
  - Usernames and email addresses
  - Software versions (OS, Office, PDF creators)
  - Internal network paths and server names
  - Printer names and shared resources
  - Hidden revision history and comments
  - GPS coordinates (from images)

INSTALLATION#

# Windows only (GUI application, .NET)
# Download from: https://github.com/ElevenPaths/FOCA/releases

# Alternative: run on Linux via Wine or VM

# Dependencies
  - .NET Framework 4.7.2+
  - SQL Server Express (LocalDB, auto-installed)

BASIC WORKFLOW#

1. Create new project (File > New Project)
2. Enter target domain (e.g., example.com)
3. Configure search engines (Google, Bing, DuckDuckGo)
4. Select file types to search for
5. Click "Search All" to find documents
6. Download discovered documents
7. Extract metadata from downloaded files
8. Analyze results in the metadata panel

SUPPORTED FILE TYPES#

Type         Extensions            Metadata Extracted
----         ----------            ------------------
Office       .doc, .docx           Author, company, template, revision
             .xls, .xlsx           Printer, paths, last saved by
             .ppt, .pptx           Software version, comments
PDF          .pdf                  Creator, producer, author, dates
OpenOffice   .odt, .ods, .odp     Author, generator, statistics
Images       .jpg, .png, .gif     EXIF (GPS, camera, software)
SVG          .svg                  Editor, embedded paths
InDesign     .indd                 Author, fonts, links

SEARCH ENGINES#

# FOCA uses search engines to discover documents
# Supported: Google, Bing, DuckDuckGo, Exalead

# Search dorks used internally:
site:example.com filetype:pdf
site:example.com filetype:docx
site:example.com filetype:xlsx
site:example.com filetype:pptx

MANUAL GOOGLE DORKS FOR DOCUMENT DISCOVERY#

# Use these to manually find documents if FOCA search is slow

site:target.com filetype:pdf
site:target.com filetype:doc OR filetype:docx
site:target.com filetype:xls OR filetype:xlsx
site:target.com filetype:ppt OR filetype:pptx
site:target.com filetype:odt OR filetype:ods
site:target.com filetype:rtf
site:target.com filetype:xml
site:target.com filetype:svg

# With keywords
site:target.com filetype:pdf "confidential"
site:target.com filetype:xlsx "internal"
site:target.com filetype:docx "password"
site:target.com filetype:pdf "draft"

METADATA EXTRACTED#

# User Information
  - Author / Creator name
  - Last Modified By
  - Company name
  - Email addresses (sometimes embedded)

# Software Information
  - Operating system (from paths)
  - Office version (e.g., Microsoft Office 16.0)
  - PDF creator (Acrobat, LibreOffice, LaTeX, etc.)
  - Application version numbers

# Network Information
  - Internal file paths (\\SERVER\Share\folder\)
  - Printer names (\\PRINTER01\HP LaserJet)
  - UNC paths revealing server names
  - Mapped drive letters with server references
  - Local file paths (C:\Users\jsmith\Documents\)

# Document Information
  - Creation date
  - Last modified date
  - Revision count
  - Print date
  - Template name used

FOCA ANALYSIS TABS#

# Network tab
  - Discovered servers and hostnames from metadata
  - Internal IP ranges inferred from paths
  - DNS resolution of discovered hostnames

# Metadata tab
  - Per-document breakdown of all extracted metadata
  - Users, dates, software, paths, printers

# Users tab
  - Aggregated list of all discovered usernames
  - Cross-referenced across documents
  - Useful for password spraying username lists

# Software tab
  - Software versions found across documents
  - Identifies outdated/vulnerable versions

# Servers tab
  - Internal servers discovered from UNC paths
  - Printer servers

COMMAND-LINE ALTERNATIVES#

# exiftool (cross-platform, most comprehensive)
exiftool document.pdf                       # All metadata
exiftool -a -u -g1 document.pdf            # All tags, grouped
exiftool -Author -Creator -Producer doc.pdf # Specific fields
exiftool -r -ext pdf /path/to/docs/        # Recursive directory
exiftool -csv *.pdf > metadata.csv         # CSV export
exiftool -json *.docx > metadata.json      # JSON export

# Extract all metadata from multiple files
exiftool -r -ext pdf -ext docx -ext xlsx . | grep -i "author\|creator\|company\|software"

# mat2 (metadata removal tool, also useful for viewing)
mat2 --show document.pdf                    # Show metadata
mat2 document.pdf                           # Strip metadata

# pdfinfo (PDF-specific)
pdfinfo document.pdf

# olemeta (Office documents, part of oletools)
pip install oletools
olemeta document.doc
oleid document.doc                          # Identify features

# metagoofil (automated document finder + extractor)
metagoofil -d example.com -t pdf,doc,xls -o /tmp/output -l 100

EXIFTOOL QUICK REFERENCE#

# Most common metadata fields
exiftool -Author file                       # Document author
exiftool -Creator file                      # Creator application
exiftool -Producer file                     # PDF producer
exiftool -Company file                      # Company name
exiftool -LastModifiedBy file               # Last editor
exiftool -CreateDate file                   # Creation date
exiftool -ModifyDate file                   # Last modified
exiftool -Software file                     # Software used
exiftool -Title file                        # Document title
exiftool -Subject file                      # Subject/description
exiftool -Keywords file                     # Keywords
exiftool -GPSPosition file                  # GPS coordinates (images)

# Batch operations
exiftool -Author -Creator *.pdf             # All PDFs in directory
exiftool -r -Author -csv *.docx > users.csv # Recursive CSV export

USING METADATA FOR PENTESTING#

# 1. Username enumeration
#    Extract authors → build username list for password spraying
#    Common formats: jsmith, john.smith, j.smith, smithj

# 2. Software version identification
#    Old Office/Acrobat versions → client-side exploits
#    PDF creators reveal server-side tools

# 3. Internal network mapping
#    UNC paths → server names, share names
#    File paths → drive mappings, directory structure

# 4. Email address discovery
#    Some docs embed email in author/properties
#    Cross-reference with other OSINT sources

# 5. Template analysis
#    Template paths reveal internal infrastructure
#    Custom templates may be downloadable

AUTOMATING METADATA COLLECTION#

#!/bin/bash
# Download all PDFs from a target and extract metadata
mkdir -p /tmp/foca_output
cd /tmp/foca_output

# Download documents (using wget or curl)
wget -r -l1 -A "*.pdf,*.docx,*.xlsx,*.pptx" \
     --no-parent https://target.com/ -P downloads/

# Extract all metadata
exiftool -r -csv downloads/ > all_metadata.csv

# Extract just usernames
exiftool -r -Author -LastModifiedBy downloads/ | \
    grep -v "^=" | grep -v "^-" | sort -u > users.txt

TIPS#

  - Always check multiple file types (PDF, Office, images)
  - UNC paths are gold for internal network mapping
  - Old documents often have more metadata (less scrubbing)
  - Cross-reference discovered usernames with LinkedIn
  - Government and education sites often have rich metadata
  - exiftool is the best CLI alternative to FOCA
  - metagoofil automates the download + extraction workflow
  - mat2 can strip metadata before sharing your own documents
  - Check revision history in Office docs for deleted content
  - GPS data in images can reveal physical locations