No articles match “”

Your AI gives generic advice because it doesn’t know why your priorities are what they are.

Every AI conversation starts over. The model doesn’t know your situation, your constraints, the decisions you’ve already made, or the reasoning behind them. It gives the best advice it can for a generic case — which is usually not the same as the best advice for your specific case. A private, versioned knowledge base that loads into every session fixes this at the source.

What actually happened

At the end of a long day that included a sprint planning session for a civic technology platform I run, I printed a document. It was a GitLab milestone description — one page, formatted for a meeting, listing every priority in the sprint with the reasoning behind each one. It showed which issues were P1 and why, who owned which lane and why that ownership was designed the way it was, and what each sub-sprint was actually for.

The document was accurate. Not in the “it listed the tickets” sense, but in the “it understood the strategy” sense. It knew that one P1 bug was blocking a content workflow that was blocking a partnership timeline. It knew that one collaborator runs their lane autonomously without approval loops — and why that was a deliberate structural choice, not an oversight. It captured the reasoning behind every prioritization call.

I didn’t explain any of this in the session. The model already knew.

The context layer

The reason the model already knew is that I maintain a private, versioned knowledge base — a small set of Markdown files in a git repository — that describes my actual situation. Career strategy. Current projects and their purposes. Financial position and the constraints it imposes. Decisions already made. The relationships and dynamics that shape how work gets done. Anything that would change the advice if the model had it.

This repository loads into every Claude Code session I run. The first message I send is the actual question — not the background. The model doesn’t need me to re-establish context because the context is already there, accurate, and current.

The sprint planning session that night worked because the model could connect the project’s task state (open issues, milestone assignments, current sprint) to the strategic context that explained why those tasks had the priorities they did. The PM system had the what. The knowledge base had the why. Both in the same session produced the combination.

The expensive part of working with AI isn’t the inference. It’s the context-establishment. Every session you re-explain your situation, or you don’t, and the model gives advice for a generic case. The context layer removes that cost.

What the files look like

The structure is simple. A CLAUDE.md at the root tells the model what the repository is for, what the key files are, what to update when things change, and what’s off-limits (specific credentials, identifying details). A knowledgebase.json is the fast-load artifact — a compact summary of current state that any session can read without opening the full prose documents. A knowledge/ directory holds the prose: one file per major domain, written to be read by a model giving advice to the person who wrote it.

The discipline is maintenance. The files are as useful as they are accurate. When something material changes — a decision gets made, a priority shifts, a constraint is added — the relevant file gets updated. A five-minute update after a significant decision makes every subsequent session better. Skipping it makes every subsequent session slightly wrong.

The knowledge files don’t contain everything. They contain what would change the advice. That test is strict: a lot of things feel important but wouldn’t actually change the output. Keeping the files lean is how they stay readable and maintainable.

Three patterns

The context layer shows up in three distinct ways depending on what you’re trying to do.

The exec-assistant pattern is for individuals: a personal knowledge base covering career, finances, projects, and priorities. Makes planning and career advice concrete instead of generic. Works with any AI session, no additional infrastructure.

The sprint planning pattern connects the context layer to a project management system — GitHub Issues, GitLab, Jira. The model reads both the strategic context (why things matter) and the current ticket state (what’s open). Sprint planning becomes something the model can reason about, not just summarize. The document you can print at the end of the session is the artifact: priorities listed with reasoning, ownership explicit, cross-dependencies surfaced.

The content creator pattern is a knowledge base of positions, research, and publication history. When you write, the model knows your voice, your constraints, your previous arguments. A citation layer lets you be opinionated without having to address every nuance — the structured topic overview does the epistemic work; you do the perspective work.

What it works with

The files can live anywhere Claude Code can read: a private GitHub repository, a folder on OneDrive or iCloud, a local git repository. The repository doesn’t need to be on a server. It doesn’t need to be hosted. It just needs to be somewhere your Claude Code session can reach it when it starts.

For the sprint planning pattern, connecting to a PM system requires additional setup: API tokens, project IDs, and some configuration specific to your stack. The exec-assistant and content patterns work immediately without any of that. Start with whichever pattern solves the problem you have today.

The open-source version

I’ve published the templates, pattern documentation, and prompts as an open-source repository: github.com/trentonmercer/context-layer. Everything needed to understand and implement the system is there. The templates are ready to copy. The concepts are documented. The prompts are ready to use.

The assembly — adapting the templates to your specific stack, connecting to your PM system, configuring the session setup for your workflow — is the work that requires someone who’s done it before. That’s what Mercature does.


The orchestration question isn’t which model is smartest. It’s who decides what gets handed off to which model.

A single AI model handling everything in a workflow is the most expensive way to run AI operations. The interesting architecture isn’t one model — it’s a planning layer that coordinates specialized models, dispatching each task to the right inference endpoint based on what it actually requires. The question is who controls that dispatch logic: the frontier model you’re already paying for, or a routing layer you own.

The master agent pattern

An orchestrated AI system has at least two layers. The planning layer reasons about what needs to happen, decomposes the work, and issues tool calls or sub-agent requests. The execution layer runs each discrete task — a document extraction, a code generation, a log analysis, a draft review — and returns results. The planning layer synthesizes those results and decides what to do next.

This is how Claude Code works in agentic mode, and it’s the pattern VergeIO and similar companies use when they say “Claude as a master agent.” The frontier model sits at the top of the hierarchy, plans the sequence, calls tools, and handles the synthesis. It’s a genuinely powerful architecture for complex multi-step work, because frontier models are currently better at planning and maintaining coherent context across long task chains than anything running locally.

The cost problem: the master model sees all the context. Every document fed to a sub-task, every intermediate result, every tool call output passes through a context window that’s billed at frontier pricing. For high-volume workflows — processing hundreds of support tickets, analyzing every commit in a sprint, scanning a document library — that adds up fast, and most of that volume doesn’t actually require frontier-quality reasoning.

Architecture A: Frontier-as-master with local delegation

The first hybrid approach keeps the frontier model as the orchestrator but routes routine subtasks to local or cheaper models through MCP tools or a proxy layer. Claude plans and synthesizes; a local 9B or 30B model does the bulk extraction, classification, and first-pass drafting. The frontier model only sees summaries and results, not the full raw content of every sub-task.

This is the pragmatic extension of the Claude Code pattern. It works well when the planning and synthesis genuinely require frontier capability and when the sub-task volume is manageable. The constraint is Anthropic’s harness: on a flat-rate subscription, the API isn’t available for arbitrary programmatic orchestration — you need to work through the official CLI or pay API rates. Which means the cost math depends heavily on how much of the work the frontier model actually handles versus how much stays in delegated tools.

The direction of control determines everything. In frontier-as-master, the frontier model decides what gets delegated. In local-first with escalation, your routing policy decides when the frontier gets called at all. One gives you strong planning with limited cost control. The other gives you strong cost control with limited planning leverage.

Architecture B: Local-first with frontier escalation

The second architecture inverts the default. A thin model-agnostic router — OpenRouter, LiteLLM, a self-hosted gateway — sits in front of all inference requests. Every task is evaluated by the router first: can a local model handle this at sufficient quality? If yes, it stays local at near-zero marginal cost. If the confidence is too low, the complexity is too high, or the task type explicitly requires frontier capability, the router escalates to a cloud model.

OpenRouter is useful here because it normalizes the API surface across dozens of models — Anthropic, OpenAI, Google, Mistral, open-weight models via hosted endpoints — behind a single interface. You configure routing rules by task type, quality threshold, or cost ceiling. The router handles model selection; your orchestration code doesn’t care which model runs each task as long as the output meets the contract.

The wins are significant: most routine volume stays local and free, sensitive data never leaves your infrastructure by default, and you own the routing policy rather than depending on a vendor to route correctly. The trade-off is that you have to build and operate the routing layer and define what “sufficient quality” means for each task type. The orchestration quality is only as strong as the routing logic.

The pragmatic hybrid: both together

The architecture that makes real production AI systems work at scale combines the two. A frontier model handles the planning layer — it maintains context, decomposes complex work, and synthesizes results — but the execution layer uses local-first routing underneath it. The frontier sees task specifications and results summaries, not the raw content of every sub-task. Volume work (classification, extraction, first-pass generation, log analysis) stays local. Tasks that need frontier-quality reasoning get escalated.

In practice, this looks like: Claude Code as the development session orchestrator with an MCP tool that wraps a local Ollama endpoint for routine code analysis and extraction tasks. Or a sprint automation system where the local model handles the volume work of parsing commits and summarizing changes, and the frontier model handles the architecture-level synthesis that turns those summaries into a coherent sprint brief. The local model is fast, free, and handles 80% of the task surface. The frontier handles the 20% that requires it.

Observability of the orchestration layer

A hybrid orchestration system without telemetry is a black box that claims to be saving you money. You need to know: what percentage of requests actually escalate to frontier models? What’s the cost per workflow? Which routing decisions are wrong — cases where you escalated to frontier for something a local model handled correctly, or cases where local model output failed quality checks?

The OpenTelemetry GenAI conventions apply here, but with an extra instrumentation requirement: routing decisions need to be traced alongside inference calls. Every escalation is a span. Every local model call is a span. The routing policy evaluation — which factors it considered, which threshold it crossed — should be attached to the escalation span as attributes. Without this, you can’t tune the routing policy because you can’t see what it’s doing.

The maturity signal for hybrid orchestration: can you pull up the last 24 hours and see your local vs. cloud dispatch ratio, your average cost per workflow type, and the cases where a local model result failed the quality gate? If not, you’re running a system that might be optimized or might not be — and you can’t tell which.


AI has access to your code. It doesn’t have access to why your code exists.

Your codebase tells an AI assistant what the code does. Your project management system tells it why — the issue that prompted the work, the constraint that shaped the decision, the sprint goal that scoped it. For most teams, that second layer is completely invisible to the AI. It doesn’t have to be.

The context gap

Spend a few hours pairing with an AI coding assistant on a real, long-lived codebase and the gap becomes obvious fast. The model reads the code well. It suggests reasonable implementations. But it has no idea why the architecture is shaped the way it is. It doesn’t know about the customer complaint that drove the last refactor, or the constraint that made the obvious solution wrong three years ago, or the sprint goal that means this particular ticket is blocking everything else.

That context lives in your project management system — in issues, in comments, in sprint boards, in the history of what was discussed before a decision got made. And almost none of it is accessible to AI tooling.

What a GitLab instance actually contains

A self-hosted GitLab that’s been used seriously for a few years is one of the richest sources of development context that exists. Every issue is a record of why something needed to change. Every comment thread is a record of how that decision got made. The sprint board is a real-time picture of what the team committed to and where it’s at against that commitment. The milestone history is the pattern of what gets estimated correctly and what doesn’t.

None of this is in the codebase. A developer who has been on the project for two years carries it in their head. When they leave, most of it goes with them. Making it accessible to AI tools is the same problem as keeping it accessible to humans — it requires a structured interface, not just a place to put text.

MCP as the bridge

Model Context Protocol is a standard for giving AI tools controlled, structured access to external data sources. Instead of copy-pasting issue descriptions into a chat window, an MCP server exposes live project state through a typed interface the AI can query directly — sprint state, issue details, milestone progress, commit-to-issue linkage.

The governance model matters here. Read access to issues is fine. Writing to issues, closing milestones, modifying sprint assignments — those are write operations that should require explicit confirmation, with dry-run defaults and a kill switch. The architecture enforces that split. Read is open; write requires intent.

“Connect modern LLMs to legacy systems without breaking the existing architecture.” That’s almost verbatim what the market is asking for. Most people can’t demonstrate it because they’ve only done greenfield work.

Three modes for three moments in a sprint

Sprint context isn’t a single question — it’s a different question depending on where the team is.

Kickoff: What are we committing to? The AI reads the milestone, the open issues, the backlog state, and helps the team understand what’s realistic given velocity and current capacity. It can surface blocked issues, missing estimates, and dependencies before the sprint starts.

Check-in: Are we on track? Mid-sprint, the same data tells a different story — how many days elapsed, what’s closed, what’s stalled, what needs to be cut or pulled forward. The team isn’t planning; they’re adjusting trajectory.

Retrospective: What do we carry forward? Post-sprint, the data becomes a learning artifact — what got estimated correctly, what didn’t, what patterns emerged across the milestone. The AI can draft an updated milestone description that reflects what actually happened versus what was planned.

Same underlying GitLab data. Three different frames. The value isn’t just querying issues — it’s shaping the output to the actual decision the team needs to make right now.

The broader point

Making development context accessible to AI tools is the same problem as making business data accessible to AI tools. The answer is the same: a structured interface with defined permissions, not open-ended access to everything. The difference between AI that helps a team and AI that confidently makes things worse is usually the quality of the context it’s working from.

GitLab has the context. MCP is how you give the AI a structured, auditable path to it. The governance model is what keeps that path from becoming a liability.


The most valuable comment explains why the obvious solution is wrong.

Not what the code does. Not what changed. Why the straightforward approach failed — what the business rule is, what the edge case is, what broke the last time someone tried the obvious fix. That’s the context that survives the original developer leaving. And increasingly, it’s what determines whether your AI tools are useful or dangerous.

Two kinds of comments

There are comments that describe what the code does. These are almost always redundant — the code already says what it does, and if it doesn’t, that’s a naming problem, not a comment problem.

Then there are comments that explain why the obvious approach was wrong. These are rare, and they are worth a hundred of the first kind. This validation exists because the upstream system sometimes sends null for a required field. The cache is invalidated here because the downstream service doesn’t respect ETags correctly. The timeout is set to 45 seconds because the legacy endpoint consistently takes 30–40 seconds on the first request.

Every one of those comments is a constraint that lives nowhere else in the codebase. It’s the accumulated knowledge of what the system actually does versus what it was designed to do. And it’s exactly what an AI coding assistant needs to not confidently recreate the bug you fixed six months ago.

Context decay in long-lived codebases

Every time a developer leaves a project, some context goes with them. The code stays. The tests stay. The history of why specific decisions were made — the customer conversation that explained the constraint, the incident postmortem that identified the failure mode, the week of debugging that finally isolated the root cause — that often doesn’t survive.

This has always been a problem. AI makes it more visible because an AI working from code alone will suggest the obvious solution with full confidence, not knowing that the obvious solution has a history. A developer who’s been on the project for two years knows instinctively to be skeptical. The AI doesn’t.

The engineering practices that make codebases maintainable for humans also make them usable by AI. Good commit messages. Linked issues. Sprint retrospectives that capture what was learned. These aren’t overhead — they’re infrastructure.

What your issue tracker actually is

A GitLab or Jira instance that’s been used seriously is the development team’s external memory. Every issue is a record of a decision: what the problem was, what was considered, what was chosen and why. The comments are the debate that preceded the decision. The linked commits are the implementation trail.

Most teams treat the issue tracker as a task list and close tickets as fast as possible. The teams that get the most from it — and increasingly, the most from AI tooling — treat it as a knowledge base. The issue isn’t done when the code ships. It’s done when anyone reading it six months later can understand the decision without asking the original developer.

Spec-driven development as context-first engineering

Writing the spec before the code is context-first development. The spec describes the constraint and the expected behavior. The code implements it. The test verifies it. Each layer is a different form of context, and each one is machine-readable by the AI working on the next iteration.

This is why spec-driven approaches work well in AI-assisted development: the AI has a complete context chain — what was intended, what was built, and whether it behaves correctly — rather than just the code in the middle. The constraint is documented before it becomes tribal knowledge that lives only in someone’s head.

The practical implication

None of this requires new tooling. Write commit messages that explain the why, not just the what. Link commits to the issues that prompted them. When you add a non-obvious constraint in code, add the sentence that explains it. Close issues with a summary of what was decided and why, not just a link to the merge request.

This is the unsexy version of AI readiness. Not a new model or a new platform — just a discipline around making the context that already exists actually accessible. To your future self, to your team, and increasingly, to the AI tools that will work better when they can see the full picture.


Your AI vendors all have governance controls. The problem is you have seven of them.

Every major AI provider now ships enterprise controls: SSO, audit logs, RBAC, data retention policies. That’s genuinely good. But if your organization is using Claude, ChatGPT, Copilot, Gemini, and a few internal models — each of those control planes is separate. Governance isn’t a setting you configure once. It’s an operational layer you have to maintain across all of them.

Governance isn’t one thing

The word “AI governance” gets used to mean very different things depending on who’s saying it. HR uses it to mean acceptable-use policy. Security uses it to mean data controls and vendor risk. Legal uses it to mean liability and compliance. IT uses it to mean access management and audit trails.

These are all legitimate concerns, but they’re different layers of the same problem. A useful way to separate them: policy answers what the organization is allowed to do; risk answers what could go wrong; control answers how you enforce both. And sitting beside all three is a fourth layer that’s often missing entirely — value measurement, which answers whether any of the AI investment is producing a return worth the cost and the risk.

Organizations that treat these as one problem end up with either no governance or a policy document that nobody actually enforces technically.

The vendor fragmentation problem

Anthropic, OpenAI, Microsoft, and Google have all built meaningful enterprise controls into their AI services. Claude Enterprise ships with SCIM provisioning, JIT access, RBAC, audit logs, and a Compliance API. The equivalent controls exist in ChatGPT Enterprise and Microsoft Copilot. They’re real, and they work within the scope of each product.

The problem is that each vendor’s control plane covers only that vendor. Which means the security team’s actual job is: log into the Claude admin console, log into the OpenAI admin console, log into the Microsoft admin console, check the Gemini workspace controls, and somehow produce a unified picture of what the organization’s AI usage actually looks like.

Per-vendor governance is the AI equivalent of managing your AWS, Azure, and GCP bills in three separate consoles and wondering why you can’t see your total cloud spend. The problem isn’t that each console is bad. It’s that you need a layer above them.

What a real governance layer looks like

The right mental model isn’t to replace vendor controls with something centralized — it’s to sit a governance layer above them. Each vendor remains responsible for access, retention, and audit within its own service. The governance layer gives the organization a single inventory of all AI systems in use, unified visibility into usage and cost across providers, and a control point that can enforce policy regardless of which tool is being used.

This is the architectural pattern behind tools like ServiceNow AI Control Tower, which positions itself as a cross-provider inventory and observability layer. Whether you use a platform tool or build that layer yourself, the architecture is the same: vendor controls at the bottom, unified governance above them.

The FinOps layer is usually missing

Even organizations with reasonable security governance often have no answer to the question: what are we actually getting for this? Token-based pricing, model-per-task billing, GPU compute for fine-tuning, RAG pipeline infrastructure — AI spend accumulates in ways that are genuinely hard to track against business outcomes. Cost per case processed. Cost per completed task. Cost per unit of accuracy improvement over the baseline.

FinOps practice applied to AI isn’t about cutting spend. It’s about making spend legible. Which workloads justify cloud model pricing. Which workloads could run on local infrastructure at a fraction of the cost. Which use cases are producing measurable value versus which ones are running because someone thought they should.

Most organizations don’t have this picture. Building it is foundational to any governance conversation that’s actually about the business rather than just the technology.


Local AI doesn’t replace cloud AI. It replaces the wrong use of cloud AI.

Running AI models locally has gotten genuinely practical. A quantized 9-billion-parameter model on a mid-range GPU handles most development tasks at 120 tokens per second, entirely offline, at no marginal cost. That’s not a replacement for the best frontier models — it’s a recognition that not every task requires one, and treating them as interchangeable is an expensive mistake.

The actual constraint is VRAM

When you run a model locally, the governing constraint is GPU memory. A model loads into VRAM; anything that doesn’t fit spills into system RAM, and performance falls off sharply — from 120 tokens per second down to 20-30. The practical question for every model is: can you fit the entire thing on the GPU?

Quantization is what makes this manageable. A Q4 quantization (four levels of compression from the full-precision baseline) roughly halves the model’s size while retaining most of its quality. A 9-billion-parameter model at Q4 fits on a 16GB GPU with room to spare. That’s the entry point for a fast, fully local coding assistant.

Mixture-of-experts architectures go further: only a fraction of the model’s parameters are active for any given inference, which means you can load a 35-billion-parameter model on hardware that would normally require something much smaller. The inactive layers live in system RAM; the active layers run on the GPU. You take a speed penalty, but you get access to a much more capable model than your hardware would otherwise support.

The question isn’t “local or cloud?” It’s “which workload justifies which infrastructure?” High-volume, cost-sensitive, latency-sensitive tasks belong on local hardware. Complex reasoning tasks with low volume belong in cloud models. The mistake is using the same answer for both.

The right-workload principle

Autocomplete needs to be fast. A model responding in under 100 milliseconds creates a fluid experience; a model taking 2 seconds creates friction. For autocomplete, a small fast model on local hardware outperforms a larger cloud model with network latency, regardless of raw capability differences. Speed is the metric that matters.

Agentic reasoning tasks — writing complex logic, planning a multi-step implementation, reviewing architectural decisions — benefit from more capable models and are less sensitive to latency. A task that takes 45 seconds on Claude Sonnet versus 9 minutes on a local 30B parameter model is still viable for both; the question is whether the quality difference justifies the cost difference.

Extraction, tagging, classification, and summarization at volume are the clearest case for local infrastructure. Running 500 documents through a cloud model costs real money. Running them through a local model costs electricity. The quality threshold for these tasks is lower, and local models at Q4 quantization clear it reliably.

What this means for architecture decisions

Hybrid architecture isn’t a compromise — it’s the correct answer. A well-configured local environment handles the high-volume, latency-sensitive work. Cloud models handle the tasks where quality is the primary variable and volume is low enough to make per-inference pricing reasonable. The governance question is knowing which workloads belong where, and having the infrastructure to enforce that decision rather than defaulting to whatever is most convenient.

For most organizations, the answer to “why are our AI costs higher than expected?” is that high-volume, low-complexity tasks are running through expensive frontier models because nobody made a deliberate architectural decision. The fix isn’t to use cheaper cloud models for everything. It’s to build the layer that routes tasks to the right infrastructure based on what they actually require.


Prompts are programs. Expected outputs are test assertions.

When you write a prompt and define what the expected response should look like, you’re writing a program — just not in Python. The model is the runtime. The expected output is the test. The iteration is debugging.

The structure is the same

Traditional programming uses formal syntax to specify behavior: given these inputs, produce this output. With language models, the program is written in plain English, and the expected output acts as a test assertion — not an exact match, but a semantic contract.

“Summarize this article in three bullet points.”

That prompt is a function definition. The model interprets intent rather than strict syntax. The precision is probabilistic and semantic rather than formal — but it’s still precision. You’re specifying what you want, and you’re wrong when the output doesn’t satisfy it.

Test messages as unit tests

Input: “Translate ‘Hello’ to Spanish.” Expected output: “Hola.” This is structurally identical to a unit test — except instead of asserting exact equality, you’re allowing for acceptable variation: synonyms, formatting differences, paraphrase. Your “program” becomes three things: a set of prompts, a set of test cases, and a tolerance for fuzziness. That’s behavior-driven development, written in English.

The output is part of the loop

What makes this different from traditional software is that the output itself shapes the program. You’re not just checking correctness — you’re shaping the model’s behavior by iterating on wording, structure, and edge cases. When a test fails, you don’t fix a syntax error. You debug language.

Traditional code is deterministic: same input, same output. Language models are probabilistic: same input, slightly different outputs. So your tests are less about exact matches and more about semantic correctness — constraint satisfaction, pattern alignment, soft contracts rather than strict pass/fail conditions.

Why this matters for anyone building on AI

Framing prompts as programs and expected outputs as test assertions helps explain why prompt engineering feels like coding, why evaluation datasets look like test suites, and why iteration is essential. You’re debugging language, not syntax errors.

It also points at a broader shift: programming is moving from formal instruction writing toward intent specification and example curation. The practitioners who will do this well are the ones who already think clearly about inputs, outputs, and the behavior that connects them — which is to say, the ones who already think like engineers.


Model the business, not the tools.

Most data integration projects fail slowly. They work at first — the pipeline runs, the dashboard loads — and then the source system gets upgraded, a vendor changes their export format, or a new tool replaces the old one, and everything breaks. The reason is almost always the same: the integration was built to mirror the shape of the tool instead of the shape of the business.

The distinction in practice

A field service company runs on maybe five kinds of things: customers, the equipment installed at each property, jobs that get done, calls that come in, and money that changes hands. That’s the business. The business-intelligence schema should have a table for each of those things — customers, equipment, jobs, calls, revenue — because those are the entities the owner actually thinks about when making decisions.

What those tables should not look like is the data structure of the field service management software. That software has hundreds of fields on a job record. You need maybe twelve of them. Pulling everything because you might need it someday creates noise that makes the schema harder to understand, harder to query, and much harder to maintain when the tool changes.

When the source system gets replaced, the schema shouldn’t change. Only the adapter that feeds it should change. If your integration breaks every time a vendor updates their API, you modeled the tool, not the business.

External references belong at the edge

The right way to handle the connection between your business model and your source systems is to store external identifiers as simple nullable text columns on the relevant entity. A customer row has a servicetitan_customer_id. A job row has a st_job_id. These are references to the source, not dependencies on it. The customer exists in your schema whether or not ServiceTitan has a matching record, and the external ID can be null if the data came from somewhere else.

This pattern means your schema never takes a foreign-key dependency on a third-party system. When a tool gets replaced, you write a new adapter that populates the same columns from the new source. The downstream queries, dashboards, and reports don’t know anything changed.

Start with the question, not the data

The other discipline that saves integration projects is starting with the decision you need to make rather than the data you have available. “What is our call-to-job conversion rate by technician this month?” is a question. Working backwards from that question tells you exactly which fields you need from the call data and the job data. Working forwards from the entire contents of a vendor export tells you nothing — it just gives you a very large table that nobody queries correctly.

This sounds obvious and it almost never happens. The default project shape is: export everything the tool provides, load it into a warehouse, try to figure out what questions it can answer. The better shape is: identify the five decisions the business makes every week, reverse-engineer the minimal schema that answers each one, and build adapters from the source systems to that schema.

Why it matters when AI gets involved

This distinction becomes more consequential once you’re connecting AI systems to operational data. An AI that can answer questions about your business is only as good as the business model its data represents. If the schema mirrors ServiceTitan instead of the actual business, the AI is reasoning about a tool rather than about customers and equipment and jobs. The answers will be technically correct and practically useless — and they’ll be confident about it.

The data model is the definition of what the business knows about itself. Getting that definition right before building anything on top of it is the most valuable hour of any integration project, and the one most often skipped.


AI doesn’t replace developer judgment. It amplifies whatever judgment you already have.

The debate about whether AI replaces developers is a distraction from the more important question: what kind of developer does it make you more of? The tools that make code generation cheap will increase the demand for people who understand what the generated code should actually do. Judgment scales. Mechanics don’t.

Jevons paradox and the demand explosion

When a technology makes something dramatically cheaper to produce, demand for that thing tends to explode, not collapse. When web development replaced thick-client programming in the 1990s, it didn’t eliminate developer jobs — it expanded them by orders of magnitude because the addressable market for software grew faster than the productivity gain. The people who got left behind were the ones who refused to move from VB6 to the web.

The same dynamic is already visible in AI-assisted development. Code generation is getting cheaper. The result isn’t fewer projects — it’s more projects that previously couldn’t justify the development cost. The demand for people who can architect those projects, understand their constraints, and make the judgment calls that AI won’t is growing with it.

The risk isn’t that AI takes your job. The risk is treating AI as a shortcut to skip the judgment layer — letting it generate the solution before you understand the problem — and discovering that you don’t actually have the skills to catch it when it’s confidently wrong.

Apple tested frontier models on novel benchmarks — same difficulty level, but genuinely unseen problems — and they failed at roughly one-tenth the rate of their standard benchmark scores. What the benchmarks measure is pattern recognition. What novel problems require is judgment. Those are different things.

System-level thinking is what compounds

An AI can generate a working function. It can’t reason about whether that function belongs in the architecture you’re building, whether the interface it exposes will create maintenance problems in six months, or whether the constraint it assumes is actually guaranteed by the calling code. That’s system-level thinking, and it’s what separates the developers AI makes more productive from the developers AI makes overconfident.

The discipline that builds it: write small, focused, single-purpose methods. Be explicit about input contracts and return values. Name things for what they do, not what they contain. These aren’t stylistic preferences — they’re the practices that make code machine-readable, by AI tools and by the next human who touches it. A codebase that AI can navigate correctly is a codebase with clear structure, explicit constraints, and explanatory context. That’s also just a good codebase.

The best use of AI in development is as a thinking partner, not an answer machine

There’s a meaningful difference between asking an AI to generate a solution and asking it to pressure-test your thinking. “Here’s my approach to this problem — what’s the failure mode?” produces more durable work than “write me a function that does X.” The first use builds your understanding while generating useful output. The second generates output that may or may not fit the problem you actually have.

This is the same reason experienced developers use rubber-duck debugging — the act of articulating the problem precisely is most of the work. AI tools that can respond to that articulation with informed pushback are a significant accelerant. But only if you’re the one doing the articulating, not accepting whatever gets handed back.

What this means practically

Use AI heavily for the mechanics: code generation, refactoring, writing tests for a behavior you’ve already specified, searching the codebase for relevant patterns, drafting documentation. These are tasks where the AI is fast and accurate and the cost of a wrong answer is low because you have the judgment to catch it.

Stay in the seat for the judgment calls: architecture decisions, constraint identification, the question of whether the solution to the stated problem actually solves the real problem. These are tasks where the AI is fast and confident and the cost of a wrong answer is high because it compounds into the codebase. The developers who use AI well are the ones who know which kind of task they’re working on.


You can’t govern what you can’t see.

Most AI governance programs produce a policy document and stop. They define what employees are and aren’t allowed to submit to AI tools, establish a list of approved vendors, maybe require security review for new integrations. These are all reasonable things. None of them answer the question an organization actually needs to answer: what is our AI doing in production right now, what is it costing, and is it working?

Token burn is the new CPU/memory

In traditional application infrastructure, the primary production health signals are CPU, memory, latency, and error rate. For AI workloads, there’s a fifth signal that most teams don’t track until costs become a problem: token consumption. Input tokens, output tokens, cost per request, cost per workflow, cache hit rate. These signals are the economic heartbeat of any AI system, and they behave very differently from compute costs — they scale with content complexity, not just with traffic volume.

A document-processing workflow that runs 50,000 pages through an LLM for summarization can cost $40,000 in a single month, or $8,000 with the right model selection and prompt optimization. Both produce output. The business doesn’t know which it’s running until the invoice arrives, if it doesn’t have token-level telemetry on every workflow.

An audit log is not observability. Observability is: can you see which workload is running what model, at what cost, with what quality signal, against what baseline? If those four questions don’t have answers you can pull up in two minutes, you don’t have governance — you have paperwork.

The standardization layer that actually exists

OpenTelemetry’s GenAI semantic conventions have matured into a usable standard. Spans now have defined fields for prompts and completions, tool call chains, model version and provider, input and output token counts, and agent reasoning steps. This means there’s a common instrumentation language that works across Anthropic, OpenAI, Google, and open-weight models running on local infrastructure.

The practical implication: a team that instruments their AI workloads with OpenTelemetry gets a unified trace that shows the entire request journey — user request to inference to tool call to response — in the same observability backend they already use for application monitoring. The AI call isn’t a black box anymore. It’s a span in a distributed trace, with latency, cost, and outcome attached.

What useful AI observability actually looks like

The maturity progression is relatively clear. At the basic level, you have token counts and latency per request — a proxy-level intercept that takes minutes to add and immediately surfaces cost anomalies. At the working level, you have full distributed traces, cost attribution per workload, and basic quality evaluation (a sample of production outputs reviewed for accuracy and hallucination rate). At the advanced level, you have hierarchical agent tracing, automated quality evaluation running continuously on production traffic, and drift detection that catches when a model’s output distribution shifts over time.

Most organizations aren’t at the advanced level yet. Most aren’t at the basic level either. The gap is consequential: without token-level telemetry, every AI governance conversation is a policy conversation disconnected from operational reality. You can prohibit employees from pasting sensitive data into AI tools, but you can’t verify it isn’t happening. You can set a budget for AI spending, but you can’t attribute it to specific workloads or decisions.

Where this connects to governance

Governance without telemetry is a statement of intent, not a control. The organizations that are ahead on this have figured out that the policy layer (what may we do?) and the technical layer (what are we actually doing?) have to be connected. A vendor-level audit log from each AI provider is not enough — it doesn’t span providers, doesn’t attribute cost to business decisions, and doesn’t surface quality signals. A unified telemetry layer does all three, and it’s the foundation any meaningful AI governance program has to build on before the policy discussion becomes real.