It was one of those nights.
Three terminals open. n8n workflow half-built. Claude Desktop throwing a tool call error I’d never seen before. Cold chai on the desk. And a voice note to myself that just says “Bro, why is the tool returning null AGAIN.”
I was on my fourth MCP server in two weeks. One for a CRM, one wired as a custom n8n node, one half-done boilerplate I kept copy-pasting like it owed me money. I had templates, utilities, a growing folder called mcp-graveyard/ for the servers that didn't survive contact with production.
And somewhere between retry number five and that cold chai, it hit me.
Most SaaS products are invisible to AI agents. And the agents are equally blind to them.
This post is everything I wish someone had told me. From “do I even need this?” all the way to “why is it on fire at scale?” Let’s go.
Step 0: Does Your SaaS Actually Need MCP?
Before you write a single line of MCP code, answer this. Most devs skip this step and end up building something nobody uses. Don’t be that person.
The 5-Q Diagnostic
Q1. Do your users use Zapier, Make, or n8n to automate things in your product?
If yes, your users are already telling you they want agent-style automation. MCP is the native answer to that. They’re not using Zapier because they love Zapier. They’re using it because you haven’t given them anything better yet.
Q2. Does your product hold state that changes over time, state that an external process needs to read or write?
Think: tasks, tickets, CRM records, calendar events, inventory counts, pipeline stages. If your product is a static content viewer, you probably don’t need MCP yet. But if agents need to read your data and act on it, you do.
Q3. Can your user describe a 3-step workflow involving your product in plain English?
“When a lead signs up, create a task, send a Slack message, and update the CRM stage.”
If users talk like that, your product is a node in an agentic workflow. It needs to behave like one.
Q4. Are you a dev tool, data tool, communication tool, or operations tool?
These four categories almost always need MCP. The agents your users are building touch all of these every single day.
Q5. Has any customer ever asked for an API, a webhook, or “some way to automate” your product?
That’s product-market fit signal for MCP. They want automation interfaces. MCP is the modern answer.
Scoring:
Score What It Means 4–5 Yes Ship MCP yesterday. Seriously. 2–3 Yes Build it now, at least for your top 5 use cases 0–1 Yes Wait until your user research changes. Don’t force it.
What Is MCP, Actually?
MCP (Model Context Protocol ) is Anthropic’s open standard for letting AI models interact with external tools and data in a structured, discoverable way.
Think of it like this: your existing REST API is built for humans writing code. MCP is built for AI agents making decisions. Same underlying system, completely different interface contract.
Without MCP:
Agent → “I’ll try calling POST /api/v1/tasks with some JSON and hope it works”
Result: hallucinated params, wrong endpoint, broken workflow
With MCP:
Agent → reads your tool schema
→ understands what to call, what params it needs, what comes back
→ calls it correctly, every time
The agent stops guessing. That’s the whole point.
Your REST API still serves your frontend, webhooks, and existing integrations. MCP sits alongside it as a separate interface layer, not a replacement.
The Three Primitives (Explained Like You’ll Actually Use Them)
1. Tools: What the Agent Can Do
Tools are your action endpoints. Create, update, delete, trigger, send. Every tool needs a name, a description (the model uses this to decide when to call it), and an input schema.
Here’s a prod-ready example. Notice the description tells the model when to use it — not just what it does:
{
"name": "create_task",
"description": "Creates a new task in the user's active workspace. Use this when the user wants to add a to-do, action item, follow-up, or reminder. Do NOT use this for updating an existing task — use update_task instead.",
"inputSchema": {
"type": "object",
"properties": {
"title": {
"type": "string",
"description": "Clear, concise task title. Keep under 100 chars."
},
"due_date": {
"type": "string",
"description": "Due date in ISO 8601 format. Example: 2025-06-01"
},
"priority": {
"type": "string",
"enum": ["low", "medium", "high"],
"description": "Task priority. Defaults to medium if not specified."
},
"assignee_id": {
"type": "string",
"description": "User ID to assign this task to. Optional."
},
"request_id": {
"type": "string",
"description": "Idempotency key. If provided and this request was already processed, returns the original result without creating a duplicate."
}
},
"required": ["title"]
}
}That request_id field at the bottom saved me from 47 duplicate tasks in production. More on that shortly.
2. Resources: What the Agent Can Read
Resources are your context endpoints. Current workspace tasks, project status, CRM records. The state of the world the agent needs before it starts acting.
Don’t skip resources. An agent that can do things but can’t see anything is performing surgery blindfolded.
{
"uri": "saas://workspace/{workspace_id}/tasks",
"name": "workspace_tasks",
"description": "All tasks in the given workspace, sorted by due date. Fetch this before creating or updating tasks to avoid duplicates.",
"mimeType": "application/json"
}3. Prompts: Pre-Built Agent Workflows
Think of these as slash-commands that ship from your MCP server. /weekly-summary, /project-status, /onboard-new-user. Power users love these — they're reusable agent conversation starters your server provides out of the box.
The API Infra: How It All Fits Together
Here’s the actual request flow when an agent calls your MCP tool. Draw this on a whiteboard once: it’ll save hours of confusion

Your MCP server is a translation layer between the agent world and your existing API. It does not replace your API. It makes your API legible to models that don’t read API docs.
The Swagger/OpenAPI Parallel
If you already have a Swagger/OpenAPI spec, you’re 60% of the way to an MCP server. Here’s exactly how the concepts map:
OpenAPI Endpoint MCP Equivalent
──────────────────────────────────────────────────────────
GET /tasks → Resource: list_tasks
GET /tasks/{id} → Resource: get_task
POST /tasks → Tool: create_task
PUT /tasks/{id} → Tool: update_task
DELETE /tasks/{id} → Tool: delete_task
POST /tasks/bulk → Tool: bulk_create_tasks
GET /tasks?status=open → Resource: open_tasks (pre-scoped)
GET /reports/weekly → Prompt: weekly_summary
POST /workflows/trigger → Tool: trigger_workflow
The key difference: in OpenAPI you describe the endpoint. In MCP you describe the intent, you’re writing for a model that reasons about when to call your tool, not a developer who reads docs.
Practical move: Take your existing OpenAPI spec, run it through this filter “would an agent need to call this?” and start with just those endpoints. Don’t expose everything.
Rule: Expose ≤15 tools per MCP server. If you need more, split into domain-specific servers (e.g., tasks-mcp, billing-mcp, users-mcp). An agent with 200 tools makes bad decisions.
Scaling MCP: The Things That Will Break in Production
Here’s the part nobody writes about. The happy path is easy. This is the stuff from my mcp-graveyard/.
Problem 1: Duplicate Mutations (The 47 Tasks Problem)
Agents retry. n8n retries. Claude retries when it’s not sure a tool call succeeded. If your create_task creates a task every time it's called — you get duplicates. I got 47 once. The client was not happy.
Fix: Idempotency keys, always.
def create_task(title: str, request_id: str = None) -> dict:
if request_id:
# Check if we've already processed this exact request
existing = db.get_idempotency_record(request_id)
if existing:
return existing["result"] # Return original result, no duplicate created
# Actually create the task
task = task_service.create(title=title)
if request_id:
db.save_idempotency_record(
key=request_id,
result=task.to_dict(),
ttl=86400 # 24 hours is usually enough
)
return task.to_dict()
Problem 2: Garbage Error Messages
Bad error message → agent retries immediately → your API gets hammered → everything falls over.
// BAD — agent has no idea what to do, retries instantly
{ "error": "429 Too Many Requests" }
// GOOD - agent understands, adjusts, waits
{
"error": "rate_limit_exceeded",
"message": "You've made 60 requests in the last 60 seconds. Wait 30 seconds before retrying.",
"retry_after_seconds": 30,
"limit": 60,
"window": "60s"
}
The model reads the second response and backs off intelligently. It just hammers the first one until your server dies. Write error messages for the model, not just for your logs.
Problem 3: Partial Failures
Your tool calls three internal services. Two succeed, one fails. What do you return?
Define this explicitly. Don’t return a 200 with partial data silently. Don’t throw a 500 that kills the whole workflow. Return a structured partial result with recovery hints:
{
"status": "partial_success",
"completed": [
{
"action": "create_task",
"task_id": "task_abc123",
"status": "success"
}
],
"failed": [
{
"action": "send_notification",
"status": "failed",
"reason": "Notification service timed out. The task was created successfully — you may safely retry just the notification step."
}
]
}The agent now knows exactly what happened and what it can safely retry. No guessing.
Problem 4: Long-Running Operations
Some operations take 10–30 seconds (generating reports, data syncs, bulk operations). MCP tool calls are synchronous by default. If the agent waits too long, it times out, retries, and now you have the same expensive operation running twice.
Fix: Async job pattern.
// Tool: run_report — returns immediately
// Input:
{ "report_type": "weekly_summary", "workspace_id": "ws_abc" }
// Response (immediate, no waiting):
{
"job_id": "job_xyz789",
"status": "queued",
"estimated_seconds": 15,
"poll_with": "get_job_status"
}
// Tool: get_job_status - agent polls this
// Input:
{ "job_id": "job_xyz789" }
// Response when complete:
{
"job_id": "job_xyz789",
"status": "complete",
"result": {
"tasks_completed": 24,
"tasks_overdue": 3,
"summary": "..."
}
}
// Response while still running:
{
"job_id": "job_xyz789",
"status": "running",
"progress_percent": 60,
"estimated_seconds_remaining": 6
}
Expose both tools. The agent handles the polling loop. You handle the job queue. Clean separation.
Problem 5: Stale Context in Long Workflows
Agentic workflows can run for minutes. In that time, another user may have updated a record your agent read 3 minutes ago. The agent is now acting on stale data.
Fix: Versioned resources + optimistic locking.
Every resource response should include last_updated and a version:
{
"resource": "task",
"id": "task_abc123",
"title": "Review Q3 report",
"status": "open",
"assignee_id": "user_xyz",
"last_updated": "2025-04-27T10:23:41Z",
"version": 7
}In your update tools, accept an optional expected_version. If the record changed since the agent last read it, return a conflict instead of silently overwriting:
// Tool input
{
"task_id": "task_abc123",
"status": "complete",
"expected_version": 7
}
// Response if another user updated it in the meantime
{
"error": "version_conflict",
"message": "This task was updated by another user since you last read it. Re-fetch the task using get_task before retrying your update.",
"current_version": 8,
"your_version": 7
}
Agents handle this gracefully when you give them the context to do so. They blow up when you silently overwrite.
The Mental Model That Changed Everything
After building MCPs for multiple SaaS products and wiring them into n8n, Claude Desktop, and custom agent loops — here’s the framing that made everything click:
Your REST API is built for developers. Your MCP server is built for agents. Design it like a product, not a side project.
This means:
- Tool names that read like natural language — create_task, not POST_tasks_v2
- Descriptions that explain when to use the tool, not just what it does
- Error messages written for a reader that will act on them, not just log them
- Schemas tight enough that the model can’t make bad assumptions
When I stopped treating MCP as “REST but for AI” and started treating it as a new UX layer for a new category of user, everything got cleaner. Schemas tighter. Errors more useful. The 2 AM voice notes to myself got a lot less unhinged.
Pre-Ship Checklist:
- Max 15 tools per server, split by domain if you need more
- Every tool description starts with “Use this when…”
- Idempotency key accepted on every mutating tool
- Rate limit errors include retry_after_seconds
- Partial failures return structured status + recovery hint
- Long operations return job_id immediately, expose get_job_status
- Resources include last_updated + version
- Update tools support expected_version for conflict detection
- Auth validated server-side on every single tool call
- Tested with an actual agent loop, not just curl
The Bottom Line
The agent layer is not a future thing. It’s running n8n workflows right now. It lives inside Cursor, Claude Desktop, custom copilots, internal tools your users are already building. And if your SaaS doesn’t speak MCP, agents are either ignoring you entirely or hallucinating their way through your REST API.
Ship the MCP server. Write descriptions that tell the model when to act. Handle failures explicitly. Design it like it’s a product.
Because it is.
Built MCPs till 2 AM too many times. Currently maintaining boilerplates, n8n nodes, and a folder called mcp-graveyard/ that I will never delete for sentimental reasons.
Found this useful? Clap, share, or just steal the idempotency pattern. I’ll take either.