Blog Posts
CyberSOCEval: A Surprising Lesson from the Malware Reasoning Benchmark
Exploring how input structure affects LLM reasoning in cybersecurity, using JSON, TOON, Graph, and Markdown representations.
Read ArticleInside ORACLE: How AVA Triages Alerts Like a Senior Analyst
An inside look at AVA's hypothesis-driven triage framework: how it reasons like a senior analyst across six stages, and gets sharper with every alert it sees.
Read ArticleCyber LLM Benchmark Hub: A Home for Measuring AI in Cybersecurity
A central place to compare cybersecurity LLM performance across tasks, domains, and models. The hub brings together 35 benchmarks, 78 models, and 10 categories.
Read PostThe Third Verdict
Exploring the nuances of AI agent decision-making and accountability in autonomous cyber operations.
Read Article