← All writingJune 25, 2026 · 18 min read
AI Engineering Practice · Strategy & Plan

A plan to reach AI Readiness for an AI-native consultancy —
from an AI Engineering perspective.

Turning a salad of one-off agents and tools into a standardized AI Engineering platform — one that builds from a compounding set of reusable building blocks, so every next solution ships faster, safer, and cheaper than the last.

A generalized reference. This is an anonymized, white-labeled version of a real AI-engineering strategy, shared as a playbook any boutique or mid-size AI consultancy can adapt. All firm-specific names, products, people, geographies, and figures have been removed; external research, citations, and tooling recommendations are kept intact.

01 · Executive Summary

One platform, two delivery tiers, one learning loop

The firm sells itself as an AI-native consultancy, but underneath the brand the engineering practice is inconsistent.

Every project rebuilds the same building blocks, and the best of them are buried inside individual solutions where no one can see, review, or reuse them.

The only AI Engineering solution in production for customers is a low-code, high-level agent — not the deeper engineering capability that would put the firm on the leading edge of the market and deliver on its stated mission: guiding human organizations through AI-driven change, from roadmaps to impact.

The gap between what the firm sells and what it can reliably build and maintain is the single biggest risk to the firm's 2026 plans.

The fix is not more tools or more awareness sessions — it is a shift in mentality.

At its core: a modular platform whose pieces assemble like a puzzle, plus a cognitive loop where consultants bake their judgement into it — building digital twins that multiply the firm's capabilities, fuelled by its own IP.

But the deeper change is how the firm builds. Not every solution will come from the platform — some will run on low-/no-code builders when that fits the client (over-relying on them is a named risk). What makes the practice credible is the same engineering discipline behind whichever path it chooses.

02 · Strategic Directions & Plan

Five directions, sequenced into a plan

The whole plan in one view — five directions to commit to together, and the phased roadmap that delivers them. More than tools or process, the aim is a shared shift in mentality and culture — becoming an AI-native company across the board, where every team builds and works with AI as a default way of working.

Move 1 — Platform

Own one platform; two offerings by use case

  • A two-tier model on one AI engineering platform the firm owns.
  • Low-code agents on builders (e.g. Snowflake, Copilot Studio) where the data lives.
  • Custom-engineered agents on a suite of reusable modules.
  • One visible catalog — each capability has a single best implementation.
Move 2 — The IP

The cognitive loop is the moat

  • A learning loop between human capital (judgement) and token capital (the AI the firm owns).
  • Consultants encode workflows & judgement; it improves with every use.
  • The edge isn't the model — it's the loop on top.
  • Compounds into IP no competitor can copy.
Move 3 — Talent

Create the FDE tier

  • A Forward Deployed / Applied Delivery Engineer owns low-code, simpler builds.
  • Best AI engineers stay on complex systems, not simple solutions.
  • Uses scarce senior talent where it's most valuable.
  • Answers a real talent-retention risk.
Move 4 — Sovereignty

Sovereign by design = differentiator

  • Provider-portable, EU-resident, AI-Act-ready architecture.
  • Guarantees a US-parented giant legally cannot make.
  • What European buyers increasingly demand.
Move 5 — Mentality & proof

Change the mentality, not just the toolkit

  • A shared way of looking at AI — build & improve systems, not just consume products.
  • A cross-functional team sharing building blocks across projects.
  • Dogfood the firm's own solutions — beat off-the-shelf ChatGPT.
  • What compounds is the firm's discipline in how it builds. Proof > promises.

The plan — build the foundation first; the rest are outputs

The plan front-loads one real investment: Step 1 — establishing how the firm builds AI, and the technical capacity to do it. Everything after it — the first proof, reuse, a packaged offering, and scale — is an output of that foundation, not separate work. This is what directly addresses the challenges in §3: no reuse, the credibility gap, talent risk, and the maintenance gap.

1

Step 1 · The foundation Establish the way of building & technical capacity

Set how the firm builds AI — a modular way of working, engineering standards, and one stack choice. From that choice follows the technical capacity — platform, building blocks, evals, skills — that builds the firm's expertise to support its own products and ship faster and better. This is the one real investment; the steps below are its outputs.

1aStandardize the way of building

Choose the two-tier stack — a custom stack by default, low-code / vendor platforms when a use case requires it — and set the modular engineering rules & practices.

Decide: stack + standards signed off
1bBuild the technical foundations

Put the platform's key pillars in place — observability, context management, infrastructure-as-code (IaC), and reusable infrastructure blocks.

Done: observability live; IaC & first reusable infra blocks in place
1cBuild the team's expertise

Every AI engineer learns the stack and the building blocks of the firm's AI capability — becoming experts and a cross-functional, multidisciplinary team that can step into any project.

Done: any engineer can build on the platform, across projects
↳ Building on the foundation
↳ independent — run in parallel
2

Process First agent

  • Re-platform the flagship internal agent — split into modular blocks — and dogfood the firm's own tools internally.
  • Reuse those building blocks across 1+ projects.
3

Process The FDE tier

  • Free experienced AI engineers to build the firm's stack & technical capacity.
  • Use FDEs (interns, juniors) for low/no-code implementations.
4

Output Technical capacity & agile agent building

  • What the firm gains — a reusable platform it owns (modular blocks, evals, observability, IaC) = real technical capacity.
  • What it enables — agile agent building: prototype, validate and ship fast, because common challenges are solved once and reused.
  • Proof — new agents built mostly by composition; the firm can build and maintain production-grade systems for internal products and clients.
5

Outcome Sovereignty & mentality

  • Ship like a top-tier AI studio — reusing proven tools, frameworks, and the experts in them.
  • Quickly build POCs on a stack that can be deployed at scale.
  • Products ready for certifications — AI-Act-ready and auditable.
  • Sovereign — resilient to platform changes and provider shifts in the industry.
Always-on · dogfooding · internal hackathons · evals-in-CI · the cognitive loop · compliance gates — these run throughout, not as a step.

03 · Current State

The challenges to address

1. Wasted effort & no reuse ("mono" agents)

Every AI project rebuilds the same common building blocks from scratch:

  • The result is several implementations of the same capability where there should be one best version — and the team never collaborates to build the single, top-tier solution for each block.
  • The best work stays buried inside individual solutions, invisible at the top level, so no one can see, review, or improve it.
  • Engineering collaboration suffers: feedback, code reviews and shared technical expertise are hard to build when nothing is reused or visible.

2. Low-code is a real business line carrying layered risk

Low-code agents are a genuine opportunity and should continue — but today they are the firm's only production proof to customers, and they carry four stacked risks:

Talent

Frontier FOMO & retention risk

  • AI engineers feel trapped doing low-code config — a retention risk.
  • Seniors on no-code tools build technical debt, not deep capability.
  • Consulting needs a dedicated FDE role to own low/no-code delivery.
  • Without an architecture strategy, coding agents pile up technical debt and slow teams down CIO.
Literacy

Weak business↔tech translation

  • The delivery side has low technical literacy, so business needs are hard to turn into what's actually buildable.
  • Use cases get scoped without knowing the limits of low-code tools.
  • The easy 80% works; the hard parts — error handling, complex logic, memory — need real engineering Inkeep.
Day 2

No maintenance / handoff muscle

  • The firm is not building ops, maintenance and handoff capability.
  • Partners for agentic maintenance are scarce — agents drift on prompts, models and RAG.
  • If the firm can't master a solution, it can't hand it off.
Lock-in

Lock-in & sovereignty risk

  • Low-code agents are bound to a proprietary runtime and to where the data lives.
  • Big platforms can change pricing, terms or access — moves the firm can't avoid or control.
  • A sovereignty risk: the firm's practice shouldn't depend on one provider's decisions.

3. Internal tools aren't built to compound

  • Internal tools (the flagship internal agent, the internal copilot experiments) are siloed, with no shared AI strategy.
  • No framework absorbs common challenges once, so the next build is no faster — effort doesn't compound.
  • No dogfooding discipline: the firm doesn't prove its own agents beat off-the-shelf tools.

4. The credibility gap

  • The firm markets deep AI capability while the engineering foundation is inconsistent — the brand promise and the delivery reality have diverged. This is the central commercial risk.
  • Gartner expects most early agentic implementations to miss their targets on integration, governance and talent gaps — not model quality Gartner (forward-looking).
  • The firms that win are the ones with a real platform underneath.
  • The consequence: not building something better than commercial tools for the firm's own work undermines its credibility and its literacy when selling AI services.

5. The organisation can't make time

  • By the nature of consulting, engineers are pulled across client work — it's hard to carve out time to build a shared stack, platform and technical foundations.
  • Without a shared mentality, common stack and engineering practices, there's no foundation to build on — so it's easier to keep starting from scratch, and the gap compounds.
  • Every other challenge is downstream of this — a platform that is "everyone's job in their spare time" never ships.
  • The fix (§9): a protected platform team that treats the platform as a product — the precondition for everything else.

04 · Strategic Vision & the AI Ecosystem

A full AI ecosystem — tooling, process, and a mentality shift

The goal isn't "an agent builder." It's a complete AI ecosystem — where the platform, the way of building, and a shared mentality stack into something that unblocks the challenges by nature, not by heroics.

PLATFORM + WAY OF BUILDING + MENTALITY → ONE ECOSYSTEM MENTALITY shared way of looking at AI · dogfooding · collaboration THE WAY OF BUILDING modular standards · golden path · engineering practices PLATFORM & PILLARS open-source frameworks/tools · observability · context mgmt · IaC · reusable blocks · evals · catalog BUILDS UP ↑ = an AI ecosystem that unblocks the challenges by nature
Three layers compound into one ecosystem: the platform & pillars (the foundation), the way of building (standards & the paved road), and the mentality (how the team thinks about AI). Stacked, they unblock the challenges by design — not by one-off effort.

What the ecosystem unblocks — by nature

Implicit collaboration

One framework everyone improves

Consultants and engineers improve the shared framework just by building on it — the single best version of each block wins, instead of divergent forks.

↳ Unblocks #1 (no reuse) · #5 (no time)

Dogfooding

Feedback & literacy from real use

The firm runs on its own platform; real usage feeds the cognitive loop, and sales learn the capabilities from experience, not slideware.

↳ Unblocks #3 (compounding) · #4 (credibility)

Technical expertise

Depth on one common stack

Engineers go deep on a shared stack and reusable blocks instead of scattering across one-offs — real capability compounds in the team.

↳ Unblocks #2 (talent)

Maintenance capacity

Operate & hand off cleanly

Standard observability, evals and ops mean the firm can maintain what it ships and hand it off cleanly — the Day-2 muscle, built in.

↳ Unblocks #2 (Day 2)

The firm's two forms of capital are its people (human capital) and its agents (token capital). The strategy is to make them compound each other — every engagement makes both the consultants and the platform smarter.

05 · Proposed Framework & Reference Architecture

The main pillars a modular platform needs

A modular agent platform is built from small, reusable pieces assembled like a puzzle — and it stands on a handful of shared pillars, each a capability the whole team builds on instead of rebuilding per project.

Library

Reusable building blocks

Small composable pieces — skills, prompts, tools, connectors — assembled like a puzzle, not rebuilt each time.

Catalog

A visible index

One place where the single best version of each block lives, so everyone can see, review and reuse it.

Golden path

A standard way to build

One supported scaffold to start an agent, with quality and compliance built in from the first commit.

Quality

Observability & evals

See what agents do in production, and prove quality before anything ships.

Portability

Model gateway

Swap models and providers freely — never locked to one platform (sovereignty).

Operations

Ops & maintenance

Run, monitor and hand off what the firm builds — the Day-2 muscle, built in.

Build on what exists — don't reinvent
The firm should not reinvent the wheel. Leverage proven open-source frameworks and industry tools (skills, tools, connectors, evals-as-code) — and concentrate the differentiation in the library of firm-specific blocks and the cognitive loop (§7). The two-tier stack that powers these pillars is detailed in §6.

06 · Recommended Common Stack & Tooling

How the firm splits its technical capabilities

The firm delivers through two complementary paths — meeting customers where they are, while building its own depth and IP. Two ways to deliver, one engineering practice.

Low-code builders

Vendor & low-code platforms

  • Used when a customer already runs on a vendor platform, has contracts to leverage, or needs fast delivery where the data lives.
  • Lets the firm build on and grow platform partnerships — certifications & partner status that open commercial doors: referrals, co-selling, new accounts.
  • Works within vendor restrictions on the customer's terms.
  • Owned by Applied Delivery Engineers.
Own platform

The firm's AI platform

  • The firm's own modular, self-evolving platform — an orchestrator with reusable tools and a shared framework.
  • For bespoke, production-grade solutions the firm fully owns and controls.
  • The path when the firm needs depth, portability, and IP that compounds.
  • Owned by core AI Engineers.

The in-house platform — a suggested reference architecture

CLOUD · EU REGIONS User Frontend · React (Vite) web UI · AG-UI client AG-UI live state ↔ input & approvals Backend · FastAPI — managed runtime → containers as it grows THE AGENT LangGraph — runtime · state · checkpoints LangChain — models · tools · prompts decides within limits — code controls the flow WHAT IT CAN DO Tools typed, scoped actions Skills reusable know-how MCPs shared tools — as reuse grows MODELS & DATA Model hosting models & reasoning SQL · NoSQL DB memory & conversations Vector search ← Object store knowledge & documents Langfuse — records & evaluates every step: traces · evals · costs · prompts GitHub repo · pipeline ships every tagged version
A suggested blueprint, not the only option: a FastAPI backend wraps the agent (LangGraph + LangChain), which calls typed tools, skills and MCPs over cloud models & data — with Langfuse recording & evaluating every step, and GitHub shipping each tagged version. One proven starting point to adapt. (representative reference blueprint)

Quality & evals — the one place not to cut corners

Evals are automated quality checks for AI — known-good examples with a clear pass/fail, run on every change so nothing ships that quietly got worse. It's also critical to monitor cost and latency, not just quality.

  • Measure quality, don't assume it. AI fails silently — confident but wrong — so dashboards look normal while the model is wrong Pan.
  • LLMs are black boxes. Observability (tracing) is the only way to see what an agent actually did, and to debug it Braintrust.
  • Feedback interfaces feed monitoring automatically. In-product ratings, thumbs and corrections flow straight into observability — turning real usage into new test cases and feeding the cognitive loop (§7) Langfuse.
  • Start small, from real failures. A handful of real failure cases beats a giant test suite; grade the outcome, not a rigid step-by-step path Anthropic.
  • Let AI help grade — carefully. An AI can judge open-ended answers, but only once it has been checked to agree with a human expert Husain.

07 · The Cognitive Loop / Consultant Digital Twin

The firm's IP: a digital twin built by consultants themselves

The platform unlocks this stage. Once it exists, the firm can close the cognitive loop — turning its human capital (workflows, domain knowledge, continuous-improvement (Lean) discipline, client context, taste) into "token capital" that compounds. This is where the firm gains its real edge, starting from how it operates internally: a consultant digital twin built by the consultants themselves — an evolving, dynamic system. A platform is replicable; this encoded judgement is not.

🧠 HUMAN CAPITAL judgement · relationships ingenuity · pattern recognition sets the goals · connects the dots grows MORE valuable, not less ⚙️ TOKEN CAPITAL the AI capability built & OWNED workflows + judgement, encoded improves with every use human agency directs — sets ambitious goals AI amplifies — makes expertise scalable & replicable THE LEARNING LOOP ↗ the firm's IP
The learning loop. Human capital sets the goals and direction; the AI capability the firm owns (token capital) amplifies and scales that expertise — and improves with every use. Unlike a model or tool, this asset is firm-specific and compounds with every project Nadella.

How the loop is operated, concretely

  • Encode: every consultant's interaction and feedback builds the agents — so each can deploy their own 1:1 digital twin, carrying their workflows, judgement and domain know-how.
  • Run: those blocks power both internal delivery and client agents.
  • Feed back: real traces, corrections and failures flow back into the prompts and evals — the twin sharpens.
  • Govern it: the loop only compounds if evals are broad and the firm dogfoods the same thing customers get.

08 · Service Offering & Sales Enablement

Discover the offering by building and using the platform

Sales talk about AI but don't know what to actually offer. Rather than inventing a fixed catalog up front, the firm lets the offering emerge: as it builds the platform and dogfoods it, the capabilities that prove genuinely valuable become what it sells — discovered through delivery, not guessed.

Early signals point toward low-code data agents, bespoke production agents, a Day-2 maintenance service, and sovereign / AI-Act-ready architecture — but these are directions to validate, not a menu to commit to now.

It also compounds: each customer implementation unblocks a cluster of adjacent use cases the firm can replicate and offer with little extra effort — every engagement widens both the catalog and the sellable offering.

Cost efficiency, not cost reduction

"Cost reduction" reads as headcount cuts and clashes with people-first brand values. Cost efficiency — the same people delivering more, better — is on-brand, aligns with continuous-improvement (Lean/Kaizen) principles, and is what the platform actually enables. This framing should be the standard sales line.

Dogfooding as the sales engine

If the firm can't build something its own teams prefer over off-the-shelf ChatGPT / Claude / Gemini, it cannot credibly sell AI services.

  • Forces the platform to be genuinely good.
  • Teaches sales the real capabilities from experience, not slideware.
  • The flagship internal agent — make it beat off-the-shelf tools for the firm's own work, and sales sell from lived proof.

Enablement — a practical path from awareness to mastery

Skip passive awareness sessions. Follow a practical path — awareness → understanding → mastery, where people learn AI by building with it, not by sitting through slides:

  • Hands-on labs — learn on the platform, not on theory.
  • Internal hackathons — feed winning patterns straight into the catalog.
  • Solution champions / ambassadors / evangelists — drive peer adoption; Applied Delivery Engineers anchor it.
  • Coaching over decks — enough real, hands-on training to turn people into regular users.

09 · Operating Model: How the Roles Support Each Other

Roles designed to complement — each frees the next to do its best work

The Forward Deployed Engineer profile — why and how

An FDE (Forward Deployed Engineer) is, simply, an engineer who embeds directly with a customer to build and ship their solution on the ground — the model OpenAI and Anthropic now use to deliver AI on customer infrastructure.

For such a firm, this lower-tech FDE owns low-code, hands-on client delivery and the customer relationship — unblocking delivery capacity while keeping its core AI engineers on frontier work. A clear career ladder, not a low-code dead-end:

Forward Deployed Engineerembedded low-code client delivery,owns the customer relationshipon-ramp · protects senior talent Core AI Engineerbespoke Tier-2 production agents,hardest blocks, eval rigorstays on the frontier Platform Leadowns the standard, catalog,paved road & shared spineleadership pipeline A CLEAR LADDER — NOT A LOW-CODE DEAD-END
The FDE tier gives juniors a fast on-ramp and a leadership pipeline, while keeping senior engineers on frontier work. Recommended ~1:1–1:2 FDE-to-core ratio; treat the embedded seat as a tour, not a forever seat, to manage burnout.

A probable addition — Prompt Engineer

A likely future role: a Prompt Engineer who can spin up quick POCs and handle day-to-day maintenance — tuning prompts, evals and context. Lower barrier to enter, high-leverage for fast validation and upkeep, and it frees both FDEs and core engineers from routine work.

10 · EU Compliance & AI Sovereignty as a Differentiator

Sovereign by design — an edge over giants and vendor-locked consultants

Build the platform sovereign by design — able to run independently of the providers that feed it — and the firm gets two things at once: it eliminates the risks recent events have exposed, and it opens commercial doors that vendor-locked competitors can't.

Why — what recent events proved

  • In June 2026, the US government ordered Anthropic to cut off all foreign access to its top models (Fable 5 / Mythos 5) — taking them offline worldwide overnight. Even the most advanced AI can be switched off by a single government order Time.
  • Under oath, Microsoft France could not guarantee EU data would never be handed to US authorities; the US CLOUD Act compels US providers to disclose data wherever it's stored Register CLOUD Act.
  • The ICC was pushed off Microsoft Office onto open-source tools — "digital sovereignty" is now a board-level topic across Europe heise.
  • Demand is real: Gartner sees sovereign-cloud spend up ~83% YoY in Europe Gartner.

Sovereign by design — the payoff

Eliminates risk

No single point of failure

A platform that can swap models, clouds and providers — and self-host the critical path — has no single point of geopolitical or vendor failure. EU-resident and AI-Act-logged by default.

Opens doors

Provider-independent delivery

Because it runs independent of any one provider, the firm can win regulated, public-adjacent and select international accounts that vendor-locked competitors can't serve — and offer compliance as a deliverable.

Beyond compliance — an opportunity
The firm can go further than the minimum: certify the platform as AI-Act-ready / sovereign, turning a legal obligation into a competitive advantage and a sellable badge clients can trust.

Honest caveat: the EU lags on raw AI compute, so the firm differentiates on trust, jurisdiction and agility, not raw capability — don't oversell it McKinsey.

11 · Appendix: References

The evidence base

Internal sources: the firm's internal strategy artifacts (engineering standards, prior platform plan, annual business strategy, brand guidelines). External sources, grounding the recommendation (selected):

  1. Anthropic — Building agents with the Claude Agent SDK. claude.com
  2. Anthropic — Building Effective Agents. anthropic.com
  3. Anthropic — Demystifying evals for AI agents. anthropic.com
  4. Anthropic — April 2026 dogfooding post-mortem. anthropic.com
  5. Model Context Protocol — architecture & registry. modelcontextprotocol.io
  6. anthropics/skills monorepo. github.com
  7. OpenAI Agents SDK. openai.github.io
  8. LangChain/LangGraph 1.0. langchain.com · LangGraph
  9. 12-Factor Agents (HumanLayer). github.com
  10. Backstage software templates. backstage.io · Spotify Golden Paths. atspotify.com
  11. AWS Agent Registry (preview). aws.amazon.com
  12. Nx mental model. nx.dev · uv workspaces. astral.sh
  13. Snowflake Cortex Agents. docs.snowflake.com · Dataiku Agent Hub. dataiku.com
  14. LiteLLM. github.com · Promptfoo. github.com
  15. Hamel Husain — LLM-as-judge. hamel.dev
  16. Best LLM tracing tools 2026 (Braintrust). braintrust.dev · LLM incident playbook. tianpan.co
  17. Palantir FDE. palantir.com · Pragmatic Engineer on FDEs. pragmaticengineer.com · a16z. a16z.com
  18. Team Topologies. teamtopologies.com · 2025 DORA report. cloud.google.com
  19. EU AI Act — Art. 26. artificialintelligenceact.eu · Digital Omnibus (Gibson Dunn). gibsondunn.com
  20. Microsoft cannot guarantee data (Register). theregister.com · CLOUD Act vs EU Data Act. kiteworks.com
  21. CJEU — automated decisions & explainability. insideprivacy.com · EDPB AI privacy risks. edpb.europa.eu
  22. Gartner sovereign-cloud forecast. gartner.com · McKinsey sovereign AI. mckinsey.com
  23. McKinsey Lilli. mckinsey.com · Zapier AI rollout. zapier.com · BCG AI at Work 2025. bcg.com
  24. US bars foreign access to Claude Fable 5 / Mythos 5 (Time, Jun 2026). time.com
  25. The AI productivity trap — why best engineers get slower (CIO). cio.com
  26. Langfuse — LLM evaluation 101. langfuse.com
  27. Satya Nadella on the human↔AI learning loop (X). x.com

Claims marked (directional), (provisional), (heuristic) or (forward-looking) rest on single or fast-moving sources — verify before client use. Proposed KPIs (reuse rate, time-to-first-agent) are firm-specific, not industry-standard terms.

Building an AI Engineering Practice — A Reference Strategy · generalized, anonymized & shareable.
A reference AI-engineering strategy brief for boutique & mid-size AI consultancies. Current state, recommendation, and plan are kept distinct throughout; external claims are sourced inline and in §12. All firm-specific names, products, people & figures have been removed. Proof > promises.
← More writingGet in touch →