WitFoo Precinct 6 Open SOC Cybersecurity Dataset for AI Security Agent Benchmarks
An open-source enterprise cybersecurity dataset featuring 2M+ sanitized SOC events, incident provenance graphs, MITRE ATT&CK mappings, and attack reports on Hug
While research into AI security agents for automating Security Operations Center (SOC) workflows and incident response has accelerated rapidly, realistic enterprise-scale datasets have remained scarce. Cybersecurity platform vendor WitFoo has released a large-scale, sanitized enterprise security event dataset (witfoo/precinct6-cybersecurity) collected directly from production deployments of its Precinct 6.x platform, published on Hugging Face under the Apache-2.0 license.

Image source: WitFoo / @hetmehtaa
Moving beyond simplistic log snippets or synthetic prompts asking models to explain single attack tools like Mimikatz, this dataset provides over 2 million (2,011,674) sanitized security event logs spanning 158 enterprise security products and 261 lead detection rules, complete with 13,119 incident provenance graphs and structured natural-language attack analysis reports.
2 Million Real-World SOC Events and MITRE ATT&CK Mapping
The dataset reflects the realistic distribution of enterprise production network traffic:
- Three-Tier Labeling: Over 99.4% of the 2,011,674 signals represent benign background noise, while suspicious and malicious events reflecting real attack behaviors are explicitly classified.
- Heterogeneous Telemetry Integration: Captures events generated across 158 distinct enterprise security tools and correlates them using 261 detection lead rules.
- MITRE ATT&CK Context: Each attack signal and detection lead includes mappings to MITRE ATT&CK tactics and techniques, enabling direct evaluation of an AI agent's tactical reasoning and threat attribution capabilities.
For researchers requiring massive scale, WitFoo also released witfoo/precinct6-cybersecurity-100m, aggregating over 114 million (114,421,340) labeled signals across 5 contributing organizations.
Over 13,000 Incident Provenance Graphs and Attack Reports
A core differentiator of the dataset is its structural representation of incident causality rather than isolated log lines:
- Incident Provenance Graphs: Across 13,119 incidents, the dataset provides graph topologies comprising 35,133 nodes and 634,190 edges in GraphML and streaming NDJSON formats. These provenance graphs allow researchers to visualize attack progression and feed structured causal topologies directly into graph neural networks (GNNs) or graph-aware agent frameworks.
- Natural-Language Attack Reports: Deterministic, template-generated natural-language attack summaries (
graph/attack_reports.jsonl) generated from structured incident metadata serve as evaluation baselines for LLM-based security reporting and incident briefing agents.
4-Layer PII Sanitization Pipeline and Benchmark Caveats
To safely share enterprise operational data, WitFoo applied a rigorous 4-layer de-identification pipeline combining regular expressions, format-specific parsers, machine learning named-entity recognition (ML/NER), and Claude AI verification. The dataset was fully regenerated as version 2.0.0 to resolve residual identifier concerns discovered in earlier iterations.
However, security agent benchmark authors must account for a critical methodological constraint:
- Approximate Oracle Ground Truth: All event labels and incident groupings are generated by WitFoo Precinct's automated correlation engine rather than comprehensive manual verification by independent security analysts.
- Evaluation Recommendations: Grading an agent solely against these automated labels partially measures alignment with Precinct's correlation logic. To ensure robust benchmark validity, researchers are advised to pair this dataset with a smaller, manually verified human-reviewed holdout test set.
Sources
- Hugging Face Datasets: witfoo/precinct6-cybersecurity
- Hugging Face Datasets (100M): witfoo/precinct6-cybersecurity-100m
- GitHub: witfoo/dataset-from-precinct6
- WitFoo Research: Open cybersecurity research datasets
- Het Mehta X (@hetmehtaa): Dataset release announcement