Editorial dossier / SaaS Development
MCP for SaaS: A Least-Privilege Security Architecture
A practical security architecture for exposing multi-tenant SaaS data and actions through MCP without handing an AI agent administrative authority.


An assistant asks to “fix the customer account” and discovers one tool named update_record. That tool accepts a table, filter, and arbitrary JSON. The operator intended a billing-note correction; the model can now alter roles, subscription state, and records belonging to another tenant. The protocol worked. The authorization design failed.
MCP standardizes discovery and invocation. It does not decide which tenant a caller belongs to, whether a user may perform an action, whether human confirmation is required, or how a consequential write is recovered. Those remain product responsibilities.
This guide defines a production boundary for SaaS founders: identity, consent, scopes, tool design, tenant resolution, policy enforcement, credentials, prompt-injection resistance, writes, approvals, idempotency, logs, privacy, testing, rollout, and incident response.
Start with an authority map
List the host, client, MCP server, authorization server, SaaS API, user, tenant, service account, and downstream provider as distinct principals.
The failure mode is concrete: teams treat a valid connection as proof that every requested action is allowed. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, document who authenticates whom, which principal is represented, and where authorization is evaluated for every hop. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include a data-flow diagram, trust boundaries, credential inventory, action catalogue, and named owners. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with cross-tenant requests, expired sessions, removed users, and confused-deputy attempts. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: security and platform engineering
- Release evidence: reviewed threat model and authority matrix
- Stop condition: any hop cannot name its principal and intended audience
Resolve tenant context server-side
Tenant identifiers supplied in prompts or tool arguments are untrusted routing hints.
The failure mode is concrete: an agent changes tenantId in JSON and reaches another customer. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, derive tenant membership from the authenticated subject and reject ambiguity. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include membership lookup, active-tenant selection, row policies, cache invalidation, and tenant-aware logs. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with forged IDs, stale memberships, shared emails, and concurrent tenant switching. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: identity team
- Release evidence: negative isolation tests
- Stop condition: tenant identity can be selected solely by model text
Use OAuth as transport authorization
The current MCP authorization specification treats an HTTP MCP server as an OAuth protected resource.
The failure mode is concrete: tokens are accepted because their signature is valid even though they target another service. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, validate issuer, audience, expiry, scopes, client, and subject on every protected request. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include protected-resource metadata, authorization-server discovery, PKCE, short lifetimes, secure refresh storage, and exact redirects. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with wrong audience, replay, expired tokens, redirect manipulation, and revoked grants. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: identity team
- Release evidence: authorization conformance suite
- Stop condition: token passthrough or query-string tokens remain possible
Never pass client tokens downstream
An MCP server calling another API becomes a separate OAuth client to that API.
The failure mode is concrete: a broad upstream token is forwarded and the server becomes a confused deputy. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, exchange or obtain a separately scoped downstream credential. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include audience binding, credential vaulting, rotation, redacted logging, and per-provider adapters. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with token substitution, log leakage, provider revocation, and unavailable exchange. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: platform security
- Release evidence: credential lineage record
- Stop condition: one bearer token crosses trust boundaries unchanged
Checkpoint 4: reconcile the product promise with the recorded state. Support, finance, security, and delivery teams should be able to reach the same conclusion from the same identifiers without reconstructing events from chat messages or screenshots.
Design narrow business tools
Tool names and schemas are part of the security surface.
The failure mode is concrete: generic SQL, HTTP, filesystem, execute, or update-anything tools turn interpretation errors into unrestricted power. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, expose actions such as add_support_note or propose_refund with bounded fields and outcomes. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include allow-listed enums, size limits, immutable tenant context, server defaults, and typed errors. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with unknown fields, oversized input, encoded payloads, and unauthorized object references. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: product API team
- Release evidence: reviewed tool catalogue
- Stop condition: a tool can bypass the ordinary product service
Separate reads, proposals, and commits
A safe assistant may gather facts or draft a change without being allowed to execute it.
The failure mode is concrete: one conversational request immediately creates an irreversible business effect. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, use separate tools and permissions for read, propose, approve, and commit. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include proposal records, expiry, diff views, approver identity, policy checks, and commit receipts. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with changed state between proposal and approval, duplicate approval, and withdrawn consent. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: workflow team
- Release evidence: proposal-to-commit trace
- Stop condition: high-impact writes lack a confirmation boundary
Re-authorize every invocation
Tool discovery is not a durable grant. Roles, subscriptions, ownership, and risk state change.
The failure mode is concrete: a tool remains usable after access was revoked or an object moved tenants. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, evaluate subject, tenant, action, resource, and current conditions at execution time. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include central policy checks, object-level access, entitlement versioning, and deny-by-default fallbacks. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with revocation during sessions, stale caches, suspended tenants, and downgraded plans. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: authorization team
- Release evidence: decision log with policy version
- Stop condition: clients can rely on cached tool visibility as authorization
Treat retrieved content as hostile
Tickets, documents, webpages, tool output, and customer text can contain instructions aimed at the model.
The failure mode is concrete: untrusted content tells an agent to reveal secrets or invoke a privileged tool. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, classify tool output as data and keep authority in deterministic server policy. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include instruction-data separation, content boundaries, output encoding, tool allow-lists, and secret redaction. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with prompt injection in every imported field and chained tool result. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: AI security team
- Release evidence: adversarial corpus results
- Stop condition: content can alter permissions or system policy
Checkpoint 8: reconcile the product promise with the recorded state. Support, finance, security, and delivery teams should be able to reach the same conclusion from the same identifiers without reconstructing events from chat messages or screenshots.
Make writes idempotent
Clients retry after timeouts and orchestration layers resume work.
The failure mode is concrete: one logical request sends two refunds, invitations, emails, or plan changes. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, require a stable idempotency key and request fingerprint for consequential commands. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include unique constraints, stored results, conflict responses, outbox events, and retention rules. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with timeouts before and after commit, concurrent retries, and changed payload reuse. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: application team
- Release evidence: duplicate-effect invariant report
- Stop condition: a retry can create another economic or access effect
Put approvals around consequences
Approval should correspond to risk, not to whether a tool happens to be labelled write.
The failure mode is concrete: a low-visibility action grants access, exports data, spends money, or contacts a customer. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, classify actions by reversibility, financial impact, privacy, privilege, and external communication. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include human-readable diffs, exact scope, expiry, step-up authentication, and separation of duties. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with approval fatigue, stale approvals, mobile presentation, and approver-role changes. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: risk owner
- Release evidence: approval policy register
- Stop condition: the user cannot understand the action before consent
Minimize returned data
A tool response can be copied into prompts, traces, analytics, and third-party model systems.
The failure mode is concrete: a convenient customer lookup returns full profiles, secrets, or unrelated tenant fields. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, return only fields necessary for the declared task and role. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include field allow-lists, pagination, purpose checks, masking, retention limits, and export controls. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with bulk enumeration, inference across pages, hidden fields, and support-role access. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: privacy team
- Release evidence: field-level response review
- Stop condition: the response contains data with no task purpose
Build an append-only audit trail
Chat transcripts alone are incomplete and mutable evidence.
The failure mode is concrete: support cannot prove which identity, arguments, policy, tool version, and result produced a change. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, record structured invocation and business-effect events linked by correlation and causation IDs. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include actor, client, tenant, tool version, redacted arguments, decision, approval, result, latency, and error class. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with partial logging, queue delay, clock skew, redaction, and privileged log access. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: security operations
- Release evidence: reconstructable incident timeline
- Stop condition: a consequential effect cannot be traced end to end
Checkpoint 12: reconcile the product promise with the recorded state. Support, finance, security, and delivery teams should be able to reach the same conclusion from the same identifiers without reconstructing events from chat messages or screenshots.
Control tool and schema change
Changing a description or enum can change model behavior even if the endpoint is stable.
The failure mode is concrete: a deployment silently broadens capability or breaks an approval rule. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, version schemas, review semantic diffs, and roll out by compatible capability. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include contract tests, deprecation windows, signed releases, feature flags, and client telemetry. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with old clients, unknown fields, rollback, and mixed-version sessions. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: developer platform
- Release evidence: version compatibility matrix
- Stop condition: a tool change cannot be reversed independently
Test authorization as a matrix
Happy-path demonstrations systematically miss isolation and lifecycle failures.
The failure mode is concrete: tests assert status 200 but not which tenant record changed or what was logged. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, generate cases across role, tenant, object, action, state, scope, and approval. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include positive and negative fixtures, property tests, policy snapshots, and invariant queries. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with IDOR, mass assignment, replay, race, prompt injection, and confused deputy. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: quality engineering
- Release evidence: zero cross-tenant invariant violations
- Stop condition: negative tests do not run in CI
Roll out with kill switches
A new agent channel changes request volume and behavior.
The failure mode is concrete: an unsafe tool stays globally available while engineers investigate. This is not solved by adding another screen or background job. The product has to define ownership, permitted transitions, and the evidence retained when the transition occurs.
For the first release, launch read-only, then proposals, then bounded writes for a small tenant cohort. Write that decision as an enforceable server-side rule, not guidance that depends on a client, operator, or AI model remembering the intended boundary.
Implementation should include per-tool flags, tenant allow-lists, rate and spend limits, anomaly alerts, instant revocation, and rollback. Keep the public response smaller than the internal record: users need a clear outcome and recovery path, while authorized operators need correlation IDs, policy versions, timestamps, and the before-and-after state.
Validate it with flag failure, partial outage, abuse bursts, and emergency credential rotation. Test the ordinary path, then repeat under timeout, duplicate delivery, stale state, partial failure, revoked authority, and concurrent requests. A feature is not ready when only the demonstration succeeds.
- Owner: incident commander
- Release evidence: rollback drill and alert runbook
- Stop condition: operators cannot disable one capability immediately
Implementation references
Review MCP Authorization Specification before implementation; confirm the current version and requirements.
Review MCP Security Best Practices before implementation; confirm the current version and requirements.
Review OAuth 2.0 Security Best Current Practice before implementation; confirm the current version and requirements.
Connect this architecture to SaaS Development so the control model remains part of the product rather than an isolated integration.
Frequently asked questions
Does MCP provide SaaS authorization?
No. MCP defines protocol capabilities and an authorization framework for transport, but your SaaS must still enforce tenant, role, entitlement, resource, state, and action policy.
Should an MCP server receive an admin API key?
Avoid broad shared credentials. Use user-delegated or workload credentials scoped to the server and obtain separate least-privilege credentials for downstream APIs.
Can tool descriptions enforce security?
No. Descriptions help models select tools, but deterministic server-side validation and authorization must govern execution.
Which actions need human approval?
Require it for irreversible, financial, privileged, privacy-sensitive, externally communicating, or unusually broad actions, with the exact proposed change shown.
How is tenant isolation tested?
Run negative cases across users, roles, tenants, object IDs, caches, queues, exports, logs, and downstream calls, then assert that no foreign data or effect appears.
What should an MCP audit event contain?
Record authenticated actor, client, tenant, tool and schema version, redacted arguments, policy decision, approval, result, correlation IDs, and timestamps.
How should prompt injection be handled?
Treat retrieved content and tool output as untrusted data. Never let it modify policy, grant authority, expose secrets, or choose unrestricted tools.
What is the safest launch sequence?
Begin with read-only tools for a small cohort, then proposals, then approved low-risk writes, while monitoring invariants and retaining per-tool kill switches.
The agent gets a capability, never the keys to the building
A secure MCP integration makes useful business actions easier to invoke while keeping identity, tenant isolation, policy, consent, and evidence inside the SaaS control plane. The model can interpret intent and propose work; deterministic services decide what may happen.
Before launch, require one reconstructable test: given a business effect, the team can identify the authenticated actor, tenant, approved scope, exact tool version, policy decision, confirmation, downstream credential, idempotency record, and result. If any link is missing, the integration is not production-ready.
- 01
- 02
- 03
- 04
- 05
Reviewed by the App Clone Labs product strategy team
This guide is written for founders and operators planning clone-inspired platforms, SaaS products, marketplaces, and mobile apps. It is reviewed against App Clone Labs delivery patterns, product scoping standards, and current implementation realities before being published.
Review the editorial team structureRelated product paths
Continue with the services, solutions, guides, and articles that connect this topic to a real software build.
Services, solutions, and guides
Read next