What the Model Context Protocol is, why an agent wired into it becomes far more effective, and what it takes to deploy it for real.
Brahim Bousnguar · September 2026
An LLM is a brilliant colleague locked in a room with no window, no phone and no access to your systems.
About whatever you put in front of it.
About your reference data, your clients, or how things stand today.
No actions, no writes, no effect on the real world.
That's what building an agent is really about: opening a window for it — cleanly.
Before MCP, plugging an agent into your information systems meant wiring it up by hand. Every single time.
"Who can start on the insurer project in November, within budget?"
To answer it, you have to cross-reference three systems:
Skills, seniority, branch, languages.
Who's booked, for how many days, in which month.
FTEs needed, start date, day-rate budget.
Three APIs. A human needs twenty minutes. An agent should need ten seconds.
You paste the API docs into the prompt, and hope.
You have access to the HR API. Here's how to call it:
GET /api/v2/employees?skill=<str>&level=<1-5>&agency=<str>
WARNING: the parameter is called "skill", NOT "competence".
The level is an integer. Do not put it in quotes.
GET /api/v2/staffing/load?employee_id=<id>&month=<YYYY-MM>
WARNING: strictly YYYY-MM format, or you get a silent 500.
POST /api/v2/bookings { "employeeId": ..., "missionId": ... }
WARNING: camelCase here, unlike the GETs above.
The API key is: sk-live-8f2a91c4d7e...
Answer in JSON. Never invent a parameter. Always check...
This prompt exists at pretty much every company, in pretty much this shape.
It writes competence instead of skill. The API ignores the unknown parameter
and returns 200 OK with the whole catalog. The agent thinks it filtered.
The docs eat thousands of tokens on every turn. The raw API responses eat just as many. The useful session gets shorter.
The API key travels through the model's context, the logs, the conversation history. All the security rests on "don't repeat it".
The same HR wiring gets rewritten for Copilot, for the support agent, for the POC of the client next door. Nothing carries over.
Every agent has to be wired to every tool. Integration cost is a product, not a sum.
4 × 5 = 20 integrations to write, test, secure and maintain. Add one agent: +5.
An open protocol, published in late 2024, that standardizes how an AI application plugs into a tool.
MCP describes, once and for all, how a system exposes its data and its actions, so that any agent can use them without integration code.
MCP is to AI what USB-C is to hardware. Before: one proprietary cable per device. After: one port, and everything plugs in.
It's LSP, but for agents. An editor that speaks LSP understands any language. An agent that speaks MCP understands any tool.
Technically: JSON-RPC 2.0 over stdio or HTTP. Nothing exotic, and that's deliberate.
The server is the only one holding the credentials for the system it exposes. The model never sees them.
Functions the model decides to call: search, compute, write. Each one has a name, a description and a typed input schema.
model-controlled
URI-addressable content that the application can load into the context: a file, a profile, a record.
application-controlled
Reusable, parameterized conversation starters that the user triggers explicitly.
user-controlled
In practice, 90% of the value of the servers deployed today comes from tools. They're what the rest of this talk is about.
import { McpServer } from '@modelcontextprotocol/server';
import { StdioServerTransport } from '@modelcontextprotocol/server/stdio';
import * as z from 'zod/v4';
const server = new McpServer({ name: 'staffing', version: '1.0.0' });
server.registerTool(
'rechercher_collaborateurs',
{
description: "Finds available consultants by skill, "
+ "branch and month. Returns a short, sorted list.",
inputSchema: z.strictObject({
competence: z.string().describe("Lowercase, e.g. 'react', 'kafka'"),
niveau_min: z.number().int().min(1).max(5).optional(),
mois: z.string().regex(/^\d{4}-\d{2}$/, 'Expected format: YYYY-MM')
})
},
async ({ competence, niveau_min, mois }) => {
const resultats = chercher(competence, niveau_min ?? 1, mois);
return { content: [{ type: 'text', text: formater(resultats) }] };
}
);
await server.connect(new StdioServerTransport());
That's it. This server works right now with Claude Code, VS Code + Copilot,
Cursor, and any in-house agent that speaks MCP.
Watch the SDK version: this code uses @modelcontextprotocol/server 2.x.
Most tutorials online still use @modelcontextprotocol/sdk 1.x, where the import
paths and the shape of inputSchema are different.
z.strictObject({
competence: z.string()
.describe("e.g. 'react'"),
niveau_min: z.number()
.int().min(1).max(5)
.optional(),
mois: z.string()
.regex(/^\d{4}-\d{2}$/)
})
{
"type": "object",
"properties": {
"competence": { "type": "string",
"description": "e.g. 'react'" },
"niveau_min": { "type": "integer",
"minimum": 1, "maximum": 5 },
"mois": { "type": "string",
"pattern": "^\\d{4}-\\d{2}$" }
},
"required": ["competence", "mois"],
"additionalProperties": false
}
The contract is generated, sent to the model and enforced at runtime. A made-up parameter no longer gets through: it's rejected with a message the agent can read and act on.
Same agents, same systems. The protocol sits in the middle — and the tangle disappears.
4 + 5 = 9 connections. Add one agent: +1, and it inherits every existing server.
MCP is an open standard. Who already speaks it natively today:
Claude & Claude Code · VS Code + GitHub Copilot · Cursor · Zed · JetBrains · Windsurf · and any agent built on an official SDK (TypeScript, Python, Java, C#, Kotlin, Go).
GitHub, Sentry, Figma, Atlassian, Stripe, Notion, Cloudflare, Postgres, Playwright… and above all: your own, exposing your systems.
The MCP server you write for a client project isn't a throwaway adapter for one specific tool. It's an asset that works with today's MCP client and next year's.
Four concrete mechanisms. None of them depends on a better model: they're architectural effects.
The parameter format is described in prose, in the prompt. The model has to remember it on every call, in the middle of everything else.
When it gets it wrong, the API often answers 200 OK
and ignores the unknown parameter. The error is invisible.
The format is a schema, shipped with the tool and checked by the server before anything runs.
When it gets it wrong, it receives a precise, readable error and fixes itself on the next turn.
# The model invents a parameter "nom" that doesn't exist:
→ { "name": "rechercher_collaborateurs", "arguments": { "nom": "Clara" } }
# Without strict validation — the server ignores the unknown key:
← "5 result(s): c-007 Leila Haddad… c-003 Clara Nunes… c-004 Yanis Cherif…"
⚠ looks like success, filter never applied, the agent carries on with wrong data.
# With z.strictObject() — real capture from the demo server:
← isError: true
"Input validation error: Unrecognized key: \"nom\""
✓ the agent reads the error, rereads the schema, calls again with "competence".
Same question: "can project m-101 be staffed?"
1. GET /missions/m-101 → 1 object
2. GET /employees → 8 full objects,
every skill,
every language, every day rate
3. GET /staffing/load?e=c-001 ┐
4. GET /staffing/load?e=c-002 │ 8 calls
… ┘
11. the model cross-checks by hand:
required vs rated skills,
days sold vs capacity,
adds up FTEs, checks the budget
≈ 11 turns · ≈ 9,500 tokens (estimate)
arithmetic done by the model
1. analyser_staffing_mission
{ "mission_id": "m-101" }
← Subscription portal overhaul
Regional insurer (m-101)
Starts 2026-11, 6 months, max day rate €650
Sufficient coverage:
3 FTE available for 2 required.
Best candidates:
100/100 Sophie Marchand (c-005)
20 days · all skills covered
70/100 Clara Nunes (c-003)
20 days · missing: node (0/4 required)
≈ 1 turn · ≈ 260 tokens (measured)
arithmetic done by tested code
The right column is the real output of the demo server (translated from French); the left one is my estimate of the equivalent REST path. And the gain isn't just cost: a calculation done by code is reproducible, one done by a model is not.
The API key is in the system prompt. So it's in the model's context, in the traces, in the session history, and in anything that logs those exchanges. It's also the same key for everyone.
The server holds the secret and uses it server-side. The model sees a tool name and a schema; never a credential.
An MCP server over HTTP can require an OAuth token and expose only what the signed-in user is allowed to see. A project manager's agent and a branch manager's agent call the same tool and don't get the same rows. That's not prompt engineering — it's classic access control, in the right place.
Everything the agent does on that system goes through one component. That's where the guardrails go.
Who called which tool, with which arguments, and when. A real audit trail, not a reconstruction from the model's logs.
The server bounds pagination, caps results, limits call rates. The agent can't take your systems down by accident.
The same server, read-only for one agent and writable for another. One line of config, not a fork of the code.
The rules live in the server, tested and versioned. They don't depend on how today's prompt happens to be worded.
On my own fleet of agents, restricting a server to 4 agents out of 14 took one config key. The same need, wired by hand, would have touched 14 codebases.
| API docs in the prompt | MCP tool | |
|---|---|---|
| Agent turns | ≈ 11 (estimated) | 1 (measured) |
| Tokens used | ≈ 9,500 (estimated) | ≈ 260 (measured) |
| Malformed calls | silent, undetected | rejected, with an actionable message |
| Business calculation | done by the model, not reproducible | done by code, tested |
| Credentials | in the model's context | server-side only |
| Reuse | rewritten per agent and per team | one server, every client |
| Audit | to be pieced together | built in, at the chokepoint |
The "MCP tool" column is measured on the demo server. The left one is an estimate of the equivalent REST path, not a brochure figure. The exact ratio depends on your domain — the direction does not change.
Every token spent reading a JSON dump, recalling a parameter format or redoing an addition is a token not spent reasoning about the problem. MCP doesn't make the model smarter: it stops wasting its time.
A complete, open-source MCP server you can clone and plug in within two minutes.
A deliberately familiar domain: staffing at a consulting firm. In-memory data, no database, no network calls — a demo shouldn't depend on anything.
| Tool | What it does | What it shows |
|---|---|---|
| rechercher_collaborateurs | Filters by skill, level, branch, availability | The schema replaces the docs |
| obtenir_collaborateur | The detailed profile for a single id | You pay for detail only when you ask for it |
| analyser_staffing_mission | Fit score, gaps, FTEs, budget alert | The server computes, not the model |
| reserver_collaborateur | Makes a booking, or fails explicitly | A write that never lies |
Four tools, not forty. That's a design choice — more on that in part 5.
server.registerTool(
'analyser_staffing_mission',
{
description: "Assesses which consultants can cover a project: fit "
+ "score, skill gaps, estimated cost and budget alert. "
+ "A single answer, already computed.",
inputSchema: z.strictObject({ mission_id: z.string(), candidats_max: z.number().default(3) }),
outputSchema: z.object({ etp_requis: z.number(), etp_couvert: z.number(), candidats: z.array(/* … */) })
},
async ({ mission_id, candidats_max }) => {
const mission = missions.find(m => m.id === mission_id);
if (!mission) return erreur(`Unknown mission '${mission_id}'. Available: ${ids}.`);
// The skills × workload × budget cross-check happens HERE, in tested TypeScript,
// not in the model's context.
const candidats = evaluer(mission).slice(0, candidats_max);
return {
content: [{ type: 'text', text: resume(mission, candidats) }], // for the model
structuredContent: { etp_requis: mission.etpRequis, etp_couvert, candidats } // for the code
};
}
);
content is the natural-language answer the model reads.
structuredContent is the same result as typed JSON, for the orchestrating code.
One call serves both.
Three commands, then one config entry. That's all that stands between an agent and these four tools.
git clone https://github.com/brbousnguar/mcp-staffing-demo
cd mcp-staffing-demo
npm install && npm run build
{
"mcpServers": {
"staffing": {
"command": "node",
"args": ["/absolute/path/to/mcp-staffing-demo/dist/index.js"]
}
}
}
claude mcp add staffing -- node …/dist/index.js
The same JSON block in .vscode/mcp.json.
The client SDK, in the language of your choice.
The server doesn't need to know anything about the client, nor the client about the server. That's precisely the point.
You ▸ Can we staff the insurer project that starts in November,
and if so, with whom?
Agent ▸ [calls analyser_staffing_mission { mission_id: "m-101" }]
Yes, the project can be covered: 3 FTE available for 2 required.
· Sophie Marchand (c-005) — 100/100. Every required skill is
covered, 20 free days in November, day rate €610, under the
€650 cap. She's the obvious pick.
· Clara Nunes (c-003) — 70/100. Strong on React and TypeScript,
but no Node experience, while the project asks for level 4.
Workable paired with Sophie, not on her own.
Heads-up: Marc Delaunay would be a perfect fit for the profile,
but he's booked at 100% until December.
You ▸ Book Sophie for 20 days in November.
Agent ▸ [calls reserver_collaborateur { … jours: 20 }]
Booking r-001 confirmed. 0 days left available.
The agent doesn't "know" anything about staffing. It can read a schema, call a tool, and relay a result computed elsewhere. That's enough — and it's verifiable.
What I learned running a fleet of agents plugged into home-built MCP servers, every day, for several months.
An agent had to fix an entry in one of my tracking tools. It called the update
function with a name parameter — a parameter the model had made up.
The server, being permissive, ignored the unknown key, changed nothing, and replied success: true.
Not the model. A server that accepted invalid input and returned a success that did nothing.
The demo server applies both. The switch from z.object to
z.strictObject was written while preparing this talk —
because the very first test reproduced the bug right away.
My first tools mirrored the table's columns. The result: for a single business intent, the agent chained five calls and had to stitch them together itself — with a one-in-five chance of picking the wrong field.
By exposing the two quantities the user actually tells apart, the same intent became a single call, with no judgment call left to the model. The loop vanished overnight.
"What intent is my user expressing?" — not "which row of my database am I exposing?" An MCP server is not an ORM over HTTP. It's an API designed for a reader who never asks a clarifying question.
Its name, description and schema are sent to the model on every turn. Forty tools means a context budget spent before the first question is even asked.
The more alike the tools look, the more the model hesitates. Two nearly identical search functions are worse than a single well-named one.
Expose to each agent only the servers it needs. For me, that's a config rule on the gateway; in a company, it will be a policy.
The right scope: one MCP server covers one domain, with tools you can tell apart in one sentence. If you have to explain the difference between two tools, the model won't find it either.
Because a technology presented without its limits isn't a technology, it's a brochure.
The protocol doesn't tell you what to expose, or at what granularity. That's where most of the outcome is decided, and it can't be automated.
MCP removes integration errors, not reasoning errors. A weak model with good tools is still a weak model.
An HTTP server that has to tell each user's permissions apart means OAuth, token management and security review. Neither free nor instant.
It gets deployed, monitored, updated, and it breaks. On my own fleet, the most common failure was never the model: it was a misconfigured listen address.
Every organization plugging agents into its systems either pays the gap between N×M and N+M — or pockets it.
Today, every client project that plugs AI into a client's systems writes its own wiring. That wiring gets thrown away when the project ends. That's where the loss is.
One server per recurring domain — a master-data system, a ticketing tool, an integration platform. Written once, versioned, tested.
The connector written for one client becomes the template for the next. You no longer deliver a script, you deliver a component.
"We can make your systems usable by your agents, with the audit and access control that go with it" is an offering. Not a line on a résumé.
In the coming months, clients will ask for agents plugged into their systems. The question won't be "can you do LLMs?" but "can you connect an agent to our systems without opening everything up?"
| Question | Why it comes up early |
|---|---|
| Who approves a server? | An MCP server gives an agent real power to act on a system. Review it like an exposed API, not like a script. |
| Read-only by default? | The simplest and most effective rule: write tools are the exception, granted explicitly, never the default. |
| Whose identity does the agent carry? | A shared service account is convenient and untraceable. The user's identity costs more and saves you in an audit. |
| Where do third-party servers live? | An external MCP server runs on your side, with your secrets. Treat it like a dependency: known source, pinned version, reviewed. |
| What do we log? | Tool calls are the only reliable record of what an agent actually did. Without them, you have no answers when an incident hits. |
None of these questions is specific to AI. They're API-exposure questions — asked about a consumer that will never call support.
What I recommend: six weeks, one domain, one metric. Not a committee.
| Step | Scope | Deliverable |
|---|---|---|
| W1–W2 | Pick one heavily used internal domain — staffing, ticketing or master data — and design 4 to 6 tools at the level of business intent. | Tool specification |
| W3–W4 | Write the server: strict schemas, read-only, call logging, a test suite. | Server deployed internally |
| W5 | Plug it into the clients the teams already use (Copilot, Claude Code), with a small group. | Real usage feedback |
| W6 | Measure: time on the target questions, agent turns, errors. Decide whether to expand or stop. | Decision memo, with numbers |
If after six weeks the teams aren't reopening the tool on their own, the pilot has failed, and we say so. A pilot that can't fail measures nothing.
Open, simple, already spoken by the tools your teams use. You buy nothing and you're not locked into anyone.
Fewer turns, fewer tokens, fewer silent errors, a single control point. Without changing models.
Expose intents, not tables. Strict schemas, no empty successes, a few well-named tools.
Making a company's systems usable by agents — with audit, permissions and monitoring — is exactly what a consulting firm knows how to do.
The complete demo server and the setup instructions are open source.
github.com/brbousnguar/mcp-staffing-demo
Brahim Bousnguar