You’re fifteen minutes into a hunt. Nothing’s wrong. No alerts, no IOCs, no flashing red. But something in the data is interesting. A pattern you can’t name yet. A shape that doesn’t fit. You pull the thread. Not because a playbook told you to, but because your brain won’t let you walk away from it.
That’s curiosity. And it’s the thing that separates a hunt from a query.
I’ve spent the last year building agentic systems for threat hunting. I’ve written about the vision, the architecture, the practical steps to get started. I believe in this work. But the deeper I get into building agents that hunt, the more I bump into a question I can’t engineer around:
Can a machine actually be curious?
This isn’t a hype piece. It’s not a doomer piece either. It’s a rubric. I’m going to break curiosity into its component parts, grade agentic AI against each one, and be honest about where the gaps are. Then I’ll lay out what I think needs to happen next.
Curiosity Isn’t One Thing
We talk about curiosity like it’s a single trait. “Good hunters are curious.” Sure. But that’s like saying good athletes are fast. It doesn’t tell you anything about the specific muscles involved.
When I look at what actually happens in my brain during a hunt, the real cognitive moves and not the romantic version, I see six distinct components working together. Each one is a different kind of thinking. And each one has a very different relationship with AI.
Here’s the framework.
1. Anomaly Recognition
“That’s weird.”
The entry point to every hunt. Something deviates from expected. A service account that just ran its first interactive login. A user authenticating from two countries in an hour. A DNS query that’s just a little too long.
This is pattern-breaking detection. Your brain has a model of “normal” and something just violated it.
2. Contextual Association
“That’s weird. And I’ve seen something like this before.”
You connect the anomaly to something unrelated. Maybe you read a threat intel report a few months back about a campaign that used similar DNS patterns. Maybe you remember a colleague mentioning that the finance team just migrated to a new SaaS tool. Maybe you saw a conference talk about a technique that looks like what you’re seeing.
This is cross-domain linking. Taking knowledge from one context and applying it to another, unprompted.
3. Hypothesis Generation
“That’s weird, I’ve seen this before, and I think I know what’s happening.”
You form a testable explanation. Not just “that’s weird” but “that’s weird because maybe the attacker is using DNS tunneling for exfil, which means I should see unusually high query volumes to a single domain with high entropy in the subdomains.”
This is the scientific method kicking in. Observation becomes theory becomes experiment.
4. Selective Attention
“There are forty weird things in this data. This is the one that matters.”
Every environment is noisy. There are always anomalies. The skill isn’t finding weird things. It’s knowing which weird things to care about. Prioritizing interesting over loud. Ignoring the alert storm to focus on the quiet signal.
This is the anti-alert-fatigue muscle. The thing that lets a hunter cut through noise without a playbook telling them where to look.
5. Discomfort Tolerance
“I don’t know what this is yet, and I’m going to sit with that.”
The temptation to close the loop is enormous. Your first explanation is “probably just a misconfigured cron job.” You could stop there. Most people do.
But curiosity means staying uncomfortable. Not accepting the first plausible answer. Keeping the hunt open when closing it would feel so much better.
6. Goal Suspension
“I came in looking for lateral movement, but this is something else entirely.”
You started with a hypothesis. The data says your hypothesis is wrong. But in the process of being wrong, you stumbled into something you weren’t looking for. A curious hunter doesn’t see a failed hypothesis. They see a door opening.
This is the willingness to abandon your plan when the evidence says the real story is somewhere else. Following the data, even when it invalidates the reason you sat down.
Grading AI Against the Rubric
Now the honest part. I’ve built agentic systems that do real hunting work. I’ve seen what they’re good at and where they fall apart.
Three grades:
Replicable. AI does this well today. Better than you, usually.
Simulatable. AI can approximate it with the right scaffolding. The output looks right. The mechanism underneath isn’t the one you’re using.
Human-Only. AI can’t do this yet, and more scaffolding won’t get there.
The line between replicable and simulatable is the one that matters. Replicable means the machine is doing the thing. Simulatable means it’s producing something shaped like the thing. Whether that’s good enough depends on what you needed it for.
Anomaly Recognition: Replicable
This is where AI genuinely earns its keep. Statistical deviation, baseline comparison, z-scoring outliers, frequency analysis. Machines do this better than we do. Faster, more consistently, across more data.
If your hunt starts with “find the thing that’s different,” an agent will beat you every time. It won’t get tired at hour six. It won’t forget to check the third index. It won’t round off a number because it looked “close enough.”
Give credit where it’s due. Anomaly recognition at scale is a solved problem for agentic systems. We should stop doing this manually.
I don’t say that as a hypothetical. I run a statistical baseline system in production: z-scores on a schedule, Claude reads the outliers, Slack gets the ones worth a human. It has never once gotten bored.
Contextual Association: Simulatable
LLMs can absolutely connect an anomaly to known tradecraft, if it’s in the training data or the context window. You can stuff an agent’s context with threat intel reports, MITRE mappings, internal documentation, and past hunt results. I do this. It works.
But it’s pattern matching against known associations. The agent connects dots that have already been drawn. What it can’t do is what you did up in the framework: remember a random blog post from months back and suddenly recognize the shape of it in live data. The novel lateral connection, the one nobody’s documented yet, the one that comes from living inside a problem space for years. That’s not in a context window.
We can brute-force association by stuffing more context. But more context isn’t the same as deeper understanding.
Hypothesis Generation: Simulatable
Agents can generate hypotheses. I’ve built pipelines that take anomalies and produce structured hypotheses with testable predictions. They work. Sometimes they’re good.
But the distribution skews safe. The agent will suggest the MITRE technique before it suggests the thing nobody’s published. It’ll generate the textbook hypothesis before the creative one. You can push quality up with better scaffolding and more specific prompts, but you’re always fighting the gravity of the training data. The most interesting hypotheses, the ones that break new ground, still come from humans who are willing to be wrong in public.
Selective Attention: Simulatable (With a Harness)
This is the one I’ve changed my mind about.
Agents treat everything as equally interesting unless you build prioritization around them. For a long time I read that as a ceiling and graded this component Human-Only. That was the wrong grade. I was grading the model when I should have been grading the system.
Give an agent a real enrichment layer, asset criticality and threat intel and proximity to other recent activity on the same user or host, and it prioritizes well. Not perfectly. Well. Take that layer away and it goes right back to flat. That isn’t a property of the model. It’s a property of what you built around it.
The gut sense I was describing, ignore the noise, this is the thread, decomposes further than I wanted to admit. What does this org care about. What’s already been ruled out. What changed recently. What sits next to something that matters. Those are mostly writable down. Agents still fall apart on the novel, the ambiguous, and the local business context nobody ever wrote down. That residue is real and it still needs you. But it’s a lot smaller than a whole component of curiosity, and grading it Human-Only let me off the hook for building the context layer that would have closed most of it.
Discomfort Tolerance: Human-Only
Agents are completion machines. Their entire architecture is built around resolving the next token, finishing the response, closing the loop. They want to answer. They want to be done.
You can prompt an agent to “keep looking” or “don’t conclude yet.” It’ll comply. But there’s no intrinsic drive to stay in the uncertainty. No discomfort that pushes it to dig deeper on its own. The persistence is borrowed from you. From the prompt, the workflow, the scaffolding. Take those away and the agent closes the hunt immediately.
The best hunters I know have a borderline pathological need to understand. They can’t let it go. That restlessness isn’t in any model weights.
I wrote a whole post about when to stop hunting. This is the other half of that problem. Knowing when to stop only matters if something in you doesn’t want to.
A hunt that runs on a schedule isn’t curiosity. It’s restlessness outsourced to a cron job. That’s the best we know how to build right now, and it works. But the wanting is still yours.
Goal Suspension: Simulatable (Barely)
You can build agents that pivot when evidence contradicts the hypothesis. If the data says “this isn’t lateral movement,” a well-built agent will pivot to exploring what it actually is. This works.
But there’s a massive gap between pivoting because the logic says to and getting excited because the hypothesis broke. A curious hunter doesn’t see a failed hunt. They see a surprise. Something unexpected just happened, and unexpected things are where the best findings live.
Agents don’t feel surprise. They don’t get excited. They just follow the next instruction. And that difference matters more than it sounds like it should.
The Scorecard
Component Grade Notes Anomaly Recognition Replicable AI’s strongest suit. Stop doing this manually. Contextual Association Simulatable Works for known patterns. Fails on novel connections. Hypothesis Generation Simulatable Good but safe. The distribution skews toward textbook. Selective Attention Simulatable Works with a real enrichment layer. Collapses without one. Discomfort Tolerance Human-Only Completion machines don’t sit with ambiguity. Goal Suspension Simulatable (Barely) Can pivot on logic. Can’t feel surprise.
One replicable. Four simulatable. One human-only.
AI handles the mechanical parts of curiosity: detection, matching, structured reasoning. The messy parts are still yours.
Graded September 2026. I expect to revise it.
The Road Forward
Honest doesn’t mean hopeless. The gaps are real, but they’re not permanent.
One caveat first. Everything above this line is something I’ve built and watched fail in specific ways. Everything below it is where I think the work goes next. Those are different kinds of claims. Weigh them differently.
When I first sketched these three, they were guesses. In the months since, research has landed on two of them, which I’m choosing to read as encouraging rather than as getting scooped. I’ve linked it where it exists.
The third is still wide open. I think it’s also the most important.
1. Memory That Enables Novel Association
Current agentic context windows are either session-bound or manually stuffed. Neither works for real contextual association. What we need is closer to how hunters actually accumulate knowledge: ambient, passive, temporal. An agent that continuously ingests threat research, IR reports, environmental changes, and past hunt results, and then says “this reminds me of something from an IR report two weeks ago” without you asking it to look.
That’s not retrieval. RAG finds things that look like your query. The blog post from a few months back doesn’t look like anything. It was just nearby in time, and that’s a link similarity search cannot see. Zep builds agent memory as a temporal knowledge graph for exactly that reason: when is a first-class relationship, not metadata on a chunk. And Letta’s sleep-time compute work has agents using their idle time to reorganize memory and reason over it in advance, mining for relationships nobody has asked about yet. That’s the unprompted part, and it’s the part that matters for curiosity.
I’ve tried the boring version of this. Most hunting programs don’t have memory at all, so ATHF starts there: a hunt repo, markdown, versioned, sitting next to the hunts. It fixes hunt amnesia. An agent with those files in context does recall past hunts, and that alone raised the floor on contextual association more than any prompt trick I’ve tried. What it doesn’t do is the thing this section is about. The files wait to be read. Nothing in that system wakes up and says “the thing you wrote down in March just became relevant.” That’s the gap between memory that remembers and memory that notices.
2. Confidence Scoring That Rewards Uncertainty
Right now, agents race to conclusions. The architecture rewards completion. What if we flipped that?
Build scoring that treats “I don’t know yet, and here’s what would change my mind” as a high-value output. Not a failure state. Reward the agent for naming the three pieces of telemetry that would move its confidence, instead of forcing a binary call on thin evidence. Make “I need more information” a first-class response, not the thing it says when it runs out of things to say. An agent that knows its own limits is already halfway to curious.
Turns out there’s research on this too. OpenAI published a paper on why models hallucinate, and the part that stuck with me is simple: models get graded like students taking a test. A guess might score. “I don’t know” scores zero. So they learn to guess.
Sit with that for a second, because it reframes the grade I gave above. When I called agents completion machines, I described it like a property of the architecture. It isn’t. It’s an incentive we built. Which means it’s an incentive we can un-build.
3. Divergence Metrics for Hypothesis Generation
When an agent generates five hypotheses and all five map to well-documented MITRE techniques, the agent isn’t being curious. It’s being an encyclopedia.
We need metrics for how different an agent’s hypotheses are, from each other and from the playbook. Not rewarding random guesses. Balancing “consistent with known tradecraft” against “something we haven’t considered.” The best hypotheses live in the tension between those two.
Measure it. Optimize for it. If you only optimize for accuracy, you get an agent that never guesses wrong and never finds anything new.
This is the one I haven’t found anybody working on. If that’s you, I want to hear about it.
The Rubric Is for You Too
I built this to grade machines. It grades hunters just as well.
How’s your anomaly recognition? Probably strong. That’s why you got into hunting. How’s your contextual association? Are you reading widely enough to make lateral connections, or are you stuck in the same five blogs? How’s your hypothesis generation? Are you generating creative theories, or do you default to the MITRE matrix every time?
And the hard ones: How’s your discomfort tolerance? Do you close hunts early because “good enough” feels better than “I don’t know yet”? How’s your goal suspension? When your hypothesis is wrong, do you feel frustrated, or curious?
The component where AI still scores human-only isn’t just the machine’s blind spot. It’s your development roadmap. That’s where you stay irreplaceable. Lauren made the same argument from the other direction: the work worth keeping is the work that builds the judgment to catch the machine when it’s wrong. Discomfort tolerance is that work.
The future of threat hunting isn’t human or machine. It’s humans doing the things machines can’t, the intuitive and restless and uncomfortable things, while machines handle the rest at scale. The rubric tells you where that line is today.
It’ll move. Build the systems that push it. But know where you stand.
We put “curiosity is the hunter’s greatest tool” in the Zen of Thrunting a year ago. This is what I think it actually means.
Happy thrunting.



