Title
THE STACK, NOT THE MODEL
The Trap
WE ALL PICKED A VENDOR IN 2023
- Most firms in this room have an AI strategy that amounts to: a ChatGPT/Copilot/Claude subscription and a policy document
- That was the right move in 2023. The models were similar enough, and anything was better than nothing
- Three years later, the model landscape changes every ~6 weeks — and most firms are architecturally stuck with the one they started with
What Lock-In Actually Costs
THE BILL COMES LATER
- Average platform migration project: ~$315k in losses (Swfte enterprise survey, Jan 2026) [VERIFY before Sept]
- 45% of enterprises say vendor lock-in has ALREADY blocked them from adopting better tools
- 67% of organisations want to avoid single-AI-vendor dependency — but haven't architected against it
- Builder.ai ($1.3bn, Microsoft-backed) collapsed and left clients stranded. Vendor risk is existential, not theoretical
The Core Idea
OWN THE ORCHESTRATION. RENT THE MODEL.
- The model is a component, not the architecture
- What's actually yours: your tools, your data connections, your memory, your workflows, your guardrails
- What's rented: whichever model is best THIS QUARTER
- The test: if your vendor changed pricing or got acquired tomorrow, how long to switch? If the answer is "weeks," you're locked in. If it's "edit one config line," you're sovereign
Credibility
STOIX: 18 MONTHS IN
- UK executive search boutique, C-suite placements for VC/PE-backed businesses
- 7 people. No CTO. No engineering team
- Four production AI systems running in the business today
- We're not experts. We're practitioners — and we're only part-way through
Build #1: The Inbox That Runs Itself [LIVE]
AN AI EMPLOYEE WITH ITS OWN EMAIL ADDRESS
- Dedicated agent inbox (
stoix@agentmail.to) — separate from anyone's personal mailbox - Every 30 minutes, the agent reads unread mail, actions it, replies threaded, marks read
- Real work it does: schedules social posts when we reply "yes," answers questions about our pipeline, edits drafts on instruction
- The guardrail that matters: it can read and reply, but its delete tools are removed. Human in the loop on anything irreversible
What brokeThe agent's confirmation replies initially merged unrelated scheduling decisions into one reply — "yes to A" silently approved B. Fixed with a scoping rule. Lesson: agent bugs look like helpfulness.
Build #2: The Marketing Engine [LIVE]
HUMAN-CURATED, AI-DRAFTED
- The pipeline: podcast transcripts → Researcher extracts themes → Drafter writes LinkedIn posts in each founder's voice → Thursday 9am digest email → we pick → it schedules
- Nothing auto-publishes. Ever. We curate, it drafts
- Per-partner voice models — Rory's posts sound like Rory, Neil's like Neil (from their real sent content, not a brand book)
- Output: 4–6 quality posts/week from a 7-person firm with no marketing hire
What brokeFirst version auto-generated podcast thumbnails using tiny CRM avatar photos — unconvincing likenesses. Fix: composite from real video frame screenshots instead. Lesson: AI can't fake provenance — start from real pixels.
Build #3: The CRM That Answers Questions [LIVE]
TEN YEARS OF SEARCH DATA, QUERYABLE IN ENGLISH
- Our ATS (Loxo) connected to the agent with ~200 tools — 27 permanently removed (every delete and merge function)
- We ask: "Who are the CFO candidates we placed into manufacturing who've moved in the last 18 months?" — and get an answer in seconds
- Governance by construction: the agent CAN'T delete records; the capability doesn't exist in its toolset
- This is the pattern: don't write policies about what the AI shouldn't do — remove the tools so it can't
What brokeEarly on we trusted the vendor's "connected" status. Turns out "connected" and "working" differ — a test command reported success while the toolset was silently absent. Lesson: verify by calling the tool, not by checking the status light.
Build #4: A Model on Our Own Machine [LIVE]
ORNITH: A 9B LOCAL MODEL, RUNNING ON A LAPTOP
- Runs on an ordinary Mac, fully offline, zero per-token cost
- What it's for: first-pass work on anything confidential — candidate data, client terms, personnel matters
- What it's not for: anything you'd bet the firm on. It's a junior analyst, not a partner
- The strategic point: when a frontier model is one config line away and a local model costs nothing, you route work to the right tier. That routing is only possible if you own the orchestration
What brokehonestly, very little — because we scoped it to work it could actually do. The failure mode with local models is ambition, not capability.
What Broke (The Honest Slide)
EVERY BUILD HAD A FAILURE MODE
- The inbox agent over-approved. Fixed with scoping rules
- The thumbnails faked likenesses. Fixed with real pixels
- The CRM "connection" lied. Fixed with verification-by-calling
- Two hard-won rules: (1) guardrails belong in the toolset, not in a policy doc. (2) human-curated beats fully-automated for anything client-facing — the quality floor is what you're actually selling
The Economics
WHAT THIS ACTUALLY COSTS
- The stack (agent framework): open source, free
- Model usage: routed by task — frontier models for judgment work, cheap models for drafts, local for confidential
- Typical month: low hundreds of pounds in API costs, against output that would otherwise need a marketing hire
- The real cost was time: ~18 months of founder attention, in hours per week not days
Monday Morning: The Ladder
WHERE IS YOUR FIRM, ACTUALLY?
- L0 — Theater: AI mentioned in pitch decks, used by no one
- L1 — Personal productivity: individuals with ChatGPT accounts (most of the industry is here — 88% of HR leaders report no significant AI business value yet) [VERIFY — Gartner via Pin.com 2026]
- L2 — Team workflow: shared prompts, some automation
- L3 — Organisational infrastructure: AI connected to your systems, with guardrails (this is where it starts compounding)
- L4 — Compounding operating system: every workflow feeds the next
- L5 — Self-driving: doesn't exist yet
Close
THE MODELS WILL KEEP CHANGING. THE STACK IS YOURS.
- Don't marry a model. Build the layer that lets you date all of them
- We're documenting the build publicly — follow along, steal what's useful
- Contact: rory.mcdermott@stoix.co.uk · stoix.co.uk · Chat CFO podcast