Owning The Stack

Your models, your hardware, your network

Hey everyone!

One of the most common questions we get at Arcanum is whether you should keep on subscribing to Frontier-based models, or are we ready to make the jump to local models, and how good are they these days?

Just over the past months Apple announced the M5 Ultra Mac Studio. AI cards like the H100s and H200s are becoming available. Firms like Aikido and XBOW are benchmarking and proving open-weight models can find real vulnerabilities.

This week we're gonna explore some resources that might be able to help you make this decision and understand how we think about the future of either in-house AI or the continued use of subscription APIs.

/ The Frontier Rental Trap

Here's the problem … When you build your offensive methodology on a model you don't control, you don't actually own your methodology.

Dave Kennedy (@HackingDave) said it plainly when he announced his firm was switching to invest in internal hardware for AI: "The inconsistencies on frontier models plus how deep we are going with research is a must."

If you've been building serious agent toolchains on rented models, you've lived this. The model changes overnight. A safety update you didn't ask for. A capability regression. A behavior shift that quietly breaks a workflow you spent a month tuning. Your agent that was reliably walking attack chains on Tuesday is refusing to discuss request smuggling on Wednesday because someone tightened a filter. You file a ticket into the void. You re-tune around the new behavior. Then it changes again. You wait with bated breath for approvals and the hope that the AI companies will let ethical cyber practitioners actually use their models. Sometimes you wait until you are massively disappointed, and they launch their own cyber subscription products.

That's not a foundation you build a consulting practice on. That's a foundation someone else can move whenever they want, and every time they move it, your delivery quality moves with it.

When you own the weights and the hardware, nobody pushes an update to your model at 2 am. Your methodology stays put. That stability isn't a nice-to-have for a shop that runs on repeatable, high-quality output. It's the whole game.


(Sponsor)

AI Threat Readiness Playbook

AI is changing how organizations approach security. As attackers use AI to identify vulnerabilities and accelerate exploitation, security teams need operational strategies that can keep pace.

The AI Threat Readiness Playbook provides a practical, four-pillar framework security leaders can use to strengthen their programs.

Readers will learn how to:

  • Reduce critical attack surface exposure

  • Improve zero-day response and remediation

  • Strengthen application security with AI-assisted analysis

  • Modernize detection and response using AI-driven workflows

Share the playbook with your teams to help them better prepare for the next generation of AI-enabled threats.

/ The Open Weights Caught Up

For a long time, the honest answer to "why not run local?" was "because the local models aren't good enough."

Aikido Security, working with Philippe Dourassov (@pilvar222), ran the benchmark that should end the debate. They burned 11.7 billion tokens, pitting 10 models against 32 freshly disclosed CVEs, three runs each, to see which models can actually rediscover real vulnerabilities.

DeepSeek V4 Pro, an open-weight model, found 28 of 32. That beat Opus 5, Grok 4.6, and Claude Sol. The best open-weight model on the planet for finding vulnerabilities is one you can download and run yourself. Qwen and GLM-5.3 landed right in the same tier as the frontier closed models.

TrustedSec's self-hosted research pointed in the same direction from a different angle. Their local models nailed the common offensive tasks: SQLi auth bypass, JWT manipulation, IDOR. The best performers were 27- to 31-billion-parameter models you can run on hardware that fits under a desk.

And these aren't the only two shops keeping score. Dreadnode keeps expanding DreadIndex, their offensive security benchmark of 76 tasks across ten domains, from web pentest and crypto to reversing and vuln research, with cost and refusal data baked into every result. The names showing up in their latest evals are the same ones: Deepseek V4 Pro, GLM 5.3, Qwen 3.8 Max. Three independent benchmarks, three different methodologies, one direction. When that many people measuring from different angles land in the same place, it's no longer noise.

Aikido and TrustedSec both landed the same caveat: open weights are formidable but not yet drop-in replacements for frontier without a carefully built harness.

This really lines up with our recommendation as well. We use a gamut of models for testing and have had a large amount of success. Our success is still very deeply rooted in building agents in those agent harnesses and having to be explicit and know deeply all of the specialized tips and tricks for prompt engineering around the agents and the harnesses.

/ The Hardware Excuse Is Gone

The last thing standing between a consultancy and its own AI infrastructure was always the hardware. That barrier just fell through the floor.

Apple announced the M5 Ultra Mac Studio with up to 512GB of unified memory and 4.3x the AI compute of the last generation. That is a machine that runs capable open-weight models locally and sits quietly next to a monitor. Two years ago this was a server-room project with GPU splitting, driver pain, and a cooling plan. Now it's a purchase order and an afternoon of setup. On top of this, the optimization for local models on Apple hardware keeps on getting better with the publishing of MLX optimizations.

NVIDIA is putting compute on desks too. Their DGX Spark, a little GB10 box, already ships as a personal AI computer that runs off a wall outlet. The bigger deal is the deskside DGX Station line built on the GB300 Grace Blackwell Ultra superchip, with roughly 784GB of memory and enough muscle to run models up to the trillion-parameter class locally.

DeepSeek V4 Pro, the 1.6-trillion-parameter model that just out-hunted every frontier lab at finding CVEs, currently needs a server packed with 8x H100s to run. A box that sits under a desk and runs models in that class is shipping now. The DGX Station is still a five-figure buy, not an impulse purchase, but the category now exists, and the gap between "needs a server room" and "sits next to your monitor" is closing fast.

This is the part I'm genuinely excited about, and it's the future we want us building toward. The open-source models and the hardware to run them are advancing in lockstep. The ceiling on what a small team can own and operate itself keeps rising every quarter. For the first time, a serious offensive shop can realistically own the entire stack.

At the top end, when you're pushing the deepest research at TrustedSec's scale, Dave's 8x H100 build is the tier above all of that. Different budget, same thesis: own the compute, own the capability. And if you're asking what the napkin math is at the high end, where you build something like Dave, it is a significant investment, somewhere around $300,000 to $500,000. But if you think about the trajectory of how the models keep on getting better and how research keeps pointing to newer ways to make inference faster and better on local hardware, this is an investment in your business.

The excuse of "we can't run this ourselves" is gone. What's left is whether you decide to, and how long the window stays open.

/ You're Funding the Squeeze

The big AI labs are in a capital arms race with no precedent. The top five hyperscalers have committed north of $600 billion to infrastructure in 2026, nearly double what they spent last year. Add Stargate and the rest of the buildout, and total sector spend crosses a trillion dollars. That money buys one thing above all others: compute. GPUs, and the memory that feeds them.

Buying has a consequence. AI data centers are now projected to eat up to 70% of the world's memory production. DRAM prices are on track to climb more than 400% from the start of 2024 through the end of this year. NVIDIA's own DGX Spark jumped from $3,999 to $4,699 mid-cycle, because of memory supply constraints. The exact hardware you'd buy to run your own models is getting scarcer and pricier because labs are vacuuming up the supply.

Every dollar you spend on their API funds that buildout. You are renting capability from the same companies whose buying is quietly pricing you out of independence. They would love for renting to be the permanent default and owning to be the weird exception that gets harder every quarter. The entire business model is being the utility layer everyone has to pay to touch.

We’re not telling you to never touch a frontier API. We use them at Arcanum where they earn it. I'm telling you to see the game clearly. When you rent, you're a customer funding the thing that makes your own independence more expensive. When you own, you're off that treadmill.

This is why local control over your own infrastructure is HUGE. It's a hedge against a market a handful of companies are actively working to corner. The window to get independent is open right now. The same forces that opened it are working to close it.

(Sponsor)

Everything CTI practitioners asked for, in one drop

Almost 500 CTI practitioners just told SANS what is actually working in 2026, and it is not more feeds. Get the 2026 SANS CTI Survey, an exec summary you can hand up the chain, a field-tested starter kit from Flare's Senior Cybercrime Investigator, a seat in Flare Academy (4,000+ practitioners), and Darkroom access. One download.

/ Why This Is a Consultancy Imperative

If you run an offensive security consultancy, your AI capability is about to be a core differentiator, and you cannot build a durable one on rented infrastructure.

Think about what actually makes a shop good. It's not the model. The model is a commodity, and this quarter proved it: an open-weight download beat the frontier at finding bugs. What makes a shop good is the craft wrapped around the model. Your methodology. Your context engineering. Your agent scoping. Your curated data feeds, your fuzz corpora, your internal knowledge from a thousand past engagements. That's the harness.

It only compounds if you own the ground it sits on. Every hour you invest tuning a harness against a rented model that shifts underneath you is an hour of work someone else can invalidate. Every hour you invest tuning a harness against a model you own is an asset that keeps paying out. One is renting. The other is building equity in your own practice.

Remember the caveat both Aikido and TrustedSec landed on? Open weights need a carefully built harness to reach frontier-level results. Read that as a threat if you're renting. Read it as the opportunity if you're building. The harness is where the skill lives; it's defensible, and it's the part a competitor can't copy off a pricing page. But you only truly own it when you own the stack it runs on.

The shops that start building local capability now will have a two-year head start on the ones that wait for it to be obvious. By the time it's obvious, the head start is the moat. This is a build-versus-rent moment, and we think building is the direction for anyone serious about this as a business.

Now let me be honest about where we actually are, because I'm not selling you a fantasy. For a chunk of the hardest work we do at Arcanum, we still reach for the frontier labs, and we expect to for a while yet. The very top end of reasoning still lives behind their APIs, and pretending otherwise would be posturing. Both things are true at the same time. We lean on the big labs today, and we are building hard toward owning our own stack. We're building now, before we're forced to. In plain terms, if you want to know, our workhorse is still Opus 4.8 and 4.6, and Codex models are increasingly performing very well. We have also engineered a complete set of agents that works really well on GLM 5.2, and we expect GLM 5.3 will perform even better. We have to do tightly control prompt engineering and agent design with GLM 5.2, but when that's in play, it works really, really well. We've also had to build our own harness as OpenClaw and other harnesses. We're not cutting it for all of the edge cases in the cybersecurity field. This was a huge endeavor, and we invested the engineering time up front to make sure we had something that would scale later on when we knew models would get better.

/ Outro

If you run a shop, start the local AI conversation this quarter. Not next year. Spec a machine, pick an open-weight model, put one person on building the harness, and start compounding. If you're a practitioner inside a consultancy, this is the thing to push your leadership on now, while it's still a head start and not table stakes.

Quick plug before we go: you can still grab a seat in our Red Blue Purple AI course, live next week on September 1 and 3. Two days on putting AI to work across red, blue, and purple team ops. Sign up here.

Thanks for reading, and as always, feel free to reply or hit us up on X if something in here sparked a thought.

Happy hacking 😎


Newsletter Sources: