How to Design a Production-Ready REST API
Learn how to design predictable REST APIs with validation, authorization, idempotency, pagination, caching, concurrency control, rate limiting, observability, and safe evolution.
30 min read
The API Works Locally — Then Production Happens
Imagine we build an Order API with only two endpoints:
POST /orders
GET /orders/:orderIdLocally everything works.
We send a valid request:
POST /ordersThe order is created.
We request it:
GET /orders/order-123The order comes back.
So we think:
API done.Then real users arrive.
A mobile client sends:
POST /ordersThe server creates the order.
But the network drops before the client receives the response.
The app retries.
Now we may have:
Order 1
Order 2or even:
Payment charged twiceAnother user opens the same order in two browser tabs.
Both tabs edit version 5.
Tab A saves first.
Tab B saves later and accidentally overwrites A's changes.
Another user modifies:
{
"tenantId": "another-company"
}in the request body.
Another client sends:
GET /orders?limit=100000A payment provider becomes slow.
A proxy accidentally caches one customer's order and serves it to another.
An export takes five minutes, but the HTTP request should not stay open for five minutes.
These are production API problems.
A production-ready REST API is not just:
routes that workIt is an explicit contract for:
invalid input
authentication
authorization
retries
concurrency
pagination
caching
timeouts
dependency failures
overload
observability
API evolutionWe will design one Order API around those concerns.
Start With the Public Contract
Before writing controllers, decide what the API should mean.
Our Order API may expose:
POST /orders
GET /orders/{orderId}
GET /orders?cursor=...&limit=20
PATCH /orders/{orderId}
POST /orders/{orderId}/cancellations
POST /order-exports
GET /operations/{operationId}These routes describe:
resources
and
business actionsnot internal code.
Prefer Resource-Oriented Routes
Imagine we want to cancel an order.
A weak API design might use:
POST /orders/123/doCancelThat route describes a controller function:
doCancel()instead of a resource.
A cleaner design is:
POST /orders/123/cancellationsWe are creating a cancellation request for an order.
This becomes even more useful later if cancellation becomes asynchronous.
The cancellation can have:
its own ID
status
authorization
timestamps
failure reason
idempotency behaviorFor example:
pending
→ processing
→ completed
→ failedThat is easier to evolve than an RPC-like:
/doCancelendpoint.
Use HTTP Methods Intentionally
HTTP methods already communicate useful behavior.
| Method | Common meaning |
|---|---|
GET | Read a resource |
POST | Create something or trigger processing |
PUT | Replace a complete representation |
PATCH | Partially update |
DELETE | Remove something |
GET Should Be Safe
Example:
GET /orders/order-123The client is asking to read information.
It should not unexpectedly:
charge money
cancel order
modify inventoryThe server may still update internal things such as:
logs
metrics
access timestampsbut the client's intended business state should not change.
This is what HTTP means by a:
safe methodIdempotent Does Not Mean "Same Response"
Consider:
DELETE /draft-orders/123The first request may return:
204 No ContentThe second may return:
404 Not FoundBut the intended result is still:
draft order does not existSo the operation can still be idempotent.
Idempotency means:
Repeating the same intended request should not cause additional business effects.
This becomes especially important for POST, because POST is usually not naturally idempotent.
Validation Is Not One Thing
Developers often say:
"Validate the request."But several different validations are happening.
Think of the request as moving through layers:
HTTP validation
↓
Schema validation
↓
Authentication
↓
Authorization
↓
Domain validationEach layer answers a different question.
1. HTTP Validation
HTTP validation asks:
Can we safely read this request?Examples:
Is Content-Type correct?
Is body too large?
Is JSON syntactically valid?2. Schema Validation
Schema validation asks:
Does the request have the expected shape?For example:
{
"customerId": "uuid",
"currency": "USD",
"items": []
}Questions include:
Is customerId a UUID?
Is quantity an integer?
Is currency supported?
Are there too many items?
Are unknown fields present?3. Authentication
Authentication asks:
Who is making the request?
For example:
userId = user-123
tenantId = tenant-acmeThis usually comes from:
session
JWT
API key
OAuth tokenIf the API uses short-lived JWTs and refresh sessions, the single-device logout guide explains the session state, rotation, and revocation behavior required for real logout.
4. Authorization
Authorization asks:
Is this user allowed to perform this action?
A user may be authenticated but still not allowed to:
update this order
view another tenant's order
cancel a shipped order
create exports5. Domain Validation
Domain validation asks:
Does this action make sense according to business rules?
For example:
Does product exist?
Is product purchasable?
Is address owned by this customer?
Is inventory available?
Can this order still be cancelled?A schema library cannot answer those questions alone.
Validate Content-Type First
Suppose our endpoint only accepts JSON.
The client should send:
Content-Type: application/jsonIf they send:
Content-Type: text/plainreturn:
415 Unsupported Media Typebefore trying to parse the body.
Limit Request Body Size
Imagine someone sends:
500 MB JSON bodyto a simple order endpoint.
Even if validation eventually rejects it, the server may already have:
allocated memory
buffered data
used bandwidth
spent CPUSo body limits should be enforced before full parsing.
For example:
POST /orders
maximum body size = 64 KBExample Safe JSON Reader
class PayloadTooLargeError extends Error {}
async function readJsonBody(request: Request, maximumBytes: number) {
const contentType = request.headers.get("content-type") ?? "";
if (!contentType.toLowerCase().startsWith("application/json")) {
throw new UnsupportedMediaTypeError("Expected application/json");
}
const reader = request.body?.getReader();
if (!reader) {
throw new InvalidJsonError("Request body is required");
}
const decoder = new TextDecoder();
let bytesRead = 0;
let text = "";
while (true) {
const { value, done } = await reader.read();
if (done) break;
bytesRead += value.byteLength;
if (bytesRead > maximumBytes) {
await reader.cancel();
throw new PayloadTooLargeError();
}
text += decoder.decode(value, { stream: true });
}
text += decoder.decode();
try {
return JSON.parse(text) as unknown;
} catch {
throw new InvalidJsonError("Request body is not valid JSON");
}
}The important idea is:
limit bytes while readingnot:
read everything
then check sizeConfigure the Proxy Too
Suppose:
Nginx limit = 100 MB
Node limit = 64 KBNginx might still buffer a huge request before Node rejects it.
Your:
CDN
load balancer
reverse proxy
applicationshould use compatible limits.
Validate the Request Schema
Suppose POST /orders accepts:
{
"customerId": "customer-42",
"currency": "USD",
"items": [
{
"productId": "product-7",
"quantity": 2
}
],
"shippingAddressId": "address-3"
}With Zod:
const placeOrderSchema = z
.object({
customerId: z.string().uuid(),
currency: z.literal("USD"),
items: z
.array(
z.object({
productId: z.string().uuid(),
quantity: z.number().int().min(1).max(100),
}),
)
.min(1)
.max(50),
shippingAddressId: z.string().uuid(),
})
.strict();This validates:
types
formats
array lengths
allowed values
unknown fieldsWhy .strict() Matters
Imagine the attacker sends:
{
"customerId": "...",
"currency": "USD",
"items": [],
"tenantId": "another-company",
"status": "paid",
"totalMinor": 1
}If unknown fields are silently accepted and later passed into persistence, the client may influence server-owned data.
.strict() helps reject unexpected properties.
But there is another important protection.
Never Pass Raw Request Bodies Into Your ORM
Bad:
await db.order.create({
data: requestBody,
});Now clients may attempt to set:
tenantId
status
role
totalMinor
paymentStatus
createdByThis is called:
mass assignmentor:
over-postingInstead, construct exactly what your application accepts.
Explicitly Map the Input
const command: PlaceOrderCommand = {
customerId: parsed.data.customerId,
currency: parsed.data.currency,
items: parsed.data.items.map((item) => ({
productId: item.productId,
quantity: item.quantity,
})),
shippingAddressId: parsed.data.shippingAddressId,
};Anything not explicitly copied is ignored or rejected.
Server-Owned Values Stay Server-Owned
The client should not decide:
tenant ownership
order status
product price
total
payment state
permissionsFor example:
tenantIdshould come from authenticated identity.
Price should come from:
product catalog/databasenot:
{
"price": 1
}submitted by the client.
Domain Validation Belongs in the Use Case
Zod can validate:
quantity = positive integerbut it cannot know whether:
product exists
product is active
customer can order
inventory exists
address belongs to customerThat requires application/domain logic.
For example:
HTTP
↓
schema validation
↓
PlaceOrder use case
↓
load products
↓
validate business rules
↓
create orderThis is useful because the same business logic can later be reused by:
HTTP endpoint
queue consumer
CLI
scheduled job
testsAuthentication Is Not Authorization
Suppose we know:
type Subject = {
userId: string;
tenantId: string;
permissions: string[];
};Authentication gives us the subject.
For example:
userId = user-1
tenantId = tenant-ABut authentication alone does not mean that user can modify every order.
Scope Database Queries by Tenant
Bad:
const order = await orders.findById(orderId);Then checking tenant afterward leaves more room for mistakes.
A safer pattern is:
const order = await orders.findByIdForTenant({
orderId,
tenantId: subject.tenantId,
});The query itself becomes something like:
SELECT *
FROM orders
WHERE id = $1
AND tenant_id = $2;If no row exists:
throw new ResourceNotFoundError("order");This helps prevent cross-tenant data access.
Then Apply Permission Rules
After loading the correctly scoped resource:
authorization.require(subject, "order:update", order);Authorization may depend on:
user role
resource owner
order status
tenant policy
department
locationFrontend Authorization Is Not Security
You should hide buttons the user cannot use.
For example:
User lacks order:cancel
→ hide Cancel buttonThat is good UX.
But the API must still enforce authorization.
Users can call the endpoint manually with:
curl
Postman
custom codeThe server is the security boundary.
Use One Stable Error Format
Do not return random error shapes.
Bad:
{
"message": "Bad request"
}Then elsewhere:
{
"error": true,
"reason": "Not found"
}Then elsewhere:
{
"errors": ["No permission"]
}Clients now need special handling everywhere.
Use one stable format.
Example Error Contract
type ApiProblem = {
type: string;
title: string;
status: number;
code: string;
detail?: string;
errors?: Array<{
path: string;
code: string;
}>;
requestId: string;
};Example:
HTTP/1.1 422 Unprocessable Content
Content-Type: application/problem+json{
"type": "https://api.example.com/problems/validation-failed",
"title": "The request is invalid",
"status": 422,
"code": "validation_failed",
"errors": [
{
"path": "items.0.quantity",
"code": "too_large"
}
],
"requestId": "req-01K5"
}Machine-Readable Error Codes Matter
Clients should handle:
validation_failednot parse:
"The request is invalid"Messages may change.
Stable codes should not.
For example:
if (error.code === "order_already_cancelled") {
showAlreadyCancelled();
}Useful HTTP Status Codes
A reasonable policy:
| Situation | Status |
|---|---|
| Invalid JSON | 400 |
| Unsupported content type | 415 |
| Body too large | 413 |
| Missing/invalid authentication | 401 |
| Authenticated but forbidden | 403 |
| Resource unavailable | 404 |
| Invalid request fields | 422 |
Stale If-Match | 412 |
| Domain conflict | 409 |
| Rate limited | 429 |
| Service overloaded | 503 |
Consistency is more important than endlessly debating:
400 vs 422Pick a documented policy and follow it.
Do Not Leak Internal Errors
Never return:
SQL errors
stack traces
Redis passwords
provider secrets
internal hostnames
database topologyto clients.
Instead:
client
→ stable public error
logs/traces
→ detailed internal contextLink the two using:
requestIdMake POST Safe to Retry
This is one of the most important production API behaviors.
Imagine:
Client
↓
POST /orders
Server
↓
creates order
charges payment
Response
↓
lost because network failedThe client does not know whether the request succeeded.
So it retries:
POST /ordersWithout protection, the server may create another order.
Use an Idempotency Key
Client sends:
POST /orders
Idempotency-Key: 01K5ORDER123The rule is:
one key
=
one logical operationIf the exact operation is retried, the server returns the original result.
Idempotency Flow
A robust flow looks like:
Request arrives
↓
validate request
↓
read Idempotency-Key
↓
calculate request fingerprint
↓
claim key atomically
↓
execute operation
↓
store result
↓
return responseIf the same key comes again:
same fingerprint
→ replay resultIf the same key is reused with a different request:
rejectbecause the key no longer represents the same logical operation.
Scope Idempotency Keys
Do not make:
abc123globally unique across every operation forever.
Scope it.
For example:
tenant:acme:POST:/orders:abc123Now another tenant can safely use the same client-generated key.
Hash the Semantic Request
Suppose the first request is:
{
"customerId": "c1",
"currency": "USD"
}Then the retry uses the same key but:
{
"customerId": "c2",
"currency": "USD"
}That should not silently replay the first order.
Generate a deterministic fingerprint from the validated application command.
For example:
const fingerprint = hashCanonicalJson(command);Then store:
idempotency key
+
fingerprint
+
resultStore the Completed Result
After creating the order, store:
HTTP status
response body
important headersThen a retry can replay:
201 Createdwith the same logical result.
It should not execute the use case again.
Idempotency and Payment Providers
Suppose our database is protected, but Stripe is called twice.
We still have a problem.
Use a stable downstream idempotency key too.
For example:
orderIdcan become the provider idempotency key.
Then:
API retry
→ same order operation
payment retry
→ same provider keyThis protects more than just the local database.
Complete POST /orders Flow
A production flow may look like:
export async function postOrders(request: Request): Promise<Response> {
const requestId = request.headers.get("x-request-id") ?? crypto.randomUUID();
const requestDeadline = AbortSignal.timeout(3_000);
try {
const subject = await authenticate(request, {
signal: requestDeadline,
});
authorization.require(subject, "order:create");
const rawBody = await readJsonBody(request, 64 * 1024);
const parsed = placeOrderSchema.safeParse(rawBody);
if (!parsed.success) {
throw new RequestValidationError(
parsed.error.issues.map((issue) => ({
path: issue.path.join("."),
code: issue.code,
})),
);
}
const idempotencyKey = requireIdempotencyKey(request.headers);
const command: PlaceOrderCommand = {
tenantId: subject.tenantId,
actorId: subject.userId,
customerId: parsed.data.customerId,
currency: parsed.data.currency,
shippingAddressId: parsed.data.shippingAddressId,
items: parsed.data.items.map((item) => ({
productId: item.productId,
quantity: item.quantity,
})),
};
const fingerprint = hashCanonicalJson(command);
const result = await idempotentOrderCommands.place({
scope: `tenant:${subject.tenantId}:POST:/orders`,
key: idempotencyKey,
fingerprint,
command,
signal: requestDeadline,
});
return Response.json(result.body, {
status: result.status,
headers: {
"Content-Type": "application/json",
Location: result.location,
"Idempotency-Key": idempotencyKey,
"X-Request-Id": requestId,
},
});
} catch (error) {
const problem = mapToApiProblem(error, requestId);
logger.warn({
message: "order request failed",
requestId,
status: problem.status,
code: problem.code,
error,
});
return Response.json(problem, {
status: problem.status,
headers: {
"Content-Type": "application/problem+json",
"X-Request-Id": requestId,
...(problem.status === 401
? {
"WWW-Authenticate": "Bearer",
}
: {}),
},
});
}
}Keep Business Logic Behind the Use Case
The controller should not become:
validation
pricing
inventory
payment
database
email
authorizationall mixed together.
The application use case should own:
load authoritative prices
validate customer relationships
validate address
apply order rules
reserve inventory
persist order
coordinate payment
store idempotency resultThe controller mainly owns HTTP concerns.
Prevent Lost Updates With ETags
Consider two users editing the same order.
Current order:
version = 7Both clients read version 7.
Client A updates first.
The order becomes:
version = 8Then Client B submits changes based on version 7.
Without concurrency protection, B may overwrite A.
Use an ETag
When returning the order:
ETag: "order-order-42-v7"The ETag identifies the representation version.
Client stores it.
Conditional GET With If-None-Match
Client later requests:
GET /orders/order-42
If-None-Match: "order-order-42-v7"If nothing changed:
304 Not Modified
ETag: "order-order-42-v7"No response body is necessary.
This saves bandwidth.
Conditional Update With If-Match
Client updates:
PATCH /orders/order-42
If-Match: "order-order-42-v7"The server verifies:
current version == 7?If yes:
perform updateIf current version is already 8:
412 Precondition FailedThe client edited stale data.
412 vs 409
These are different problems.
Use:
412 Precondition Failedwhen:
client's expected version
does not match current versionUse:
409 Conflictwhen:
version was current
but business state prevents actionExample:
Order version matches
but order has already shipped
Attempt:
cancel order
Result:
409 ConflictEasy rule:
412
→ stale representation
409
→ domain/business conflictDatabase Check Must Also Be Atomic
Do not only do:
SELECT version
↓
compare in controller
↓
UPDATE laterAnother request could update between those steps.
Instead the database operation itself should check the version.
For example:
UPDATE orders
SET status = $1,
version = version + 1
WHERE id = $2
AND tenant_id = $3
AND version = $4;Then inspect:
rows affectedIf zero rows were updated:
version changed
or resource unavailableThis is optimistic concurrency control.
Never Return Unlimited Collections
Bad:
GET /ordersreturning:
every order ever createdAs the dataset grows, this becomes dangerous.
Use pagination.
For stable ordering, cursor design, backward navigation, and index selection, see the pagination-at-scale guide.
Deterministic Cursor Pagination
For example:
SELECT
id,
status,
total_minor,
created_at
FROM orders
WHERE tenant_id = $1
AND (
created_at,
id
) < ($2, $3)
ORDER BY
created_at DESC,
id DESC
LIMIT 21;The ordering:
created_at DESC
+
id DESCis deterministic.
id breaks ties when two orders have the same timestamp.
Why LIMIT 21?
If page size is:
20query:
21If 21 rows return:
hasNextPage = trueReturn only the first 20.
No separate:
COUNT(*)is required just to know whether more data exists.
Cursor Response
{
"items": [],
"pageInfo": {
"nextCursor": "opaque-signed-cursor",
"hasNextPage": true
}
}The cursor should represent:
last ordering boundary
+
filters
+
tenant scope
+
cursor versionClients should treat it as an opaque string.
Cache Behavior Must Be Explicit
Caching is part of the API contract.
The caching architecture guide develops the invalidation, stampede, hot-key, and Redis-failure decisions behind that contract.
Imagine a private order response.
We may return:
Cache-Control: private, no-cache
ETag: "order-order-42-v7"
Vary: Authorizationno-cache Does Not Mean no Storage
This confuses many developers.
Cache-Control: no-cachemeans:
cache may store response
but must revalidate before reuseMeanwhile:
Cache-Control: no-storemeans:
do not store this responseat all.
Private vs Public Caches
Authenticated customer data often needs:
privatebecause shared proxies/CDNs should not reuse it across users.
Public product catalog data may use:
Cache-Control:
public,
max-age=60,
s-maxage=300,
stale-while-revalidate=30That allows shared caching.
Never Cache Private Data Under an Unsafe Key
Suppose:
GET /orders/order-42returns tenant A's order.
If the shared cache key only contains:
/orders/order-42and ignores authentication context, another user may receive that cached response.
That is a serious data leak.
Caching and authorization must be designed together.
Long Work Should Not Hold HTTP Requests Open
Suppose exporting 2 million orders takes:
5 minutesDo not do:
POST /order-exports
↓
keep request open for 5 minutes
↓
hope connection survivesUse an asynchronous operation.
Return 202 Accepted
Client requests:
POST /order-exports
Idempotency-Key: export-01K5Server creates durable work and returns:
202 Accepted
Location: /operations/operation-123
Retry-After: 5Response:
{
"operationId": "operation-123",
"status": "pending",
"statusUrl": "/operations/operation-123"
}Operation Resource
The operation can move through:
pending
↓
running
↓
succeededor:
running
↓
failedor:
pending
↓
cancelledClient checks:
GET /operations/operation-123Persist Before Returning 202
This matters.
Bad flow:
return 202
↓
then try to enqueue job
↓
process crashesClient thinks work exists.
But it disappeared.
Safer:
create durable operation
↓
commit
↓
return 202Then workers can reliably process it.
Different Endpoints Need Different Limits
Do not apply:
100 requests/minute/IPto every endpoint.
Different routes consume different resources.
For example:
POST /orders
→ tenant request rate
→ payment-provider limitGET /orders
→ tenant rate
→ max page sizePOST /order-exports
→ monthly export quota
→ concurrency limitPOST /login
→ account + network abuse limitRate limits should protect real resources.
The rate-limiting strategies guide compares the algorithms and distributed Redis enforcement choices.
Return 429 When Limited
Example:
429 Too Many Requests
Retry-After: 10Response:
{
"error": {
"code": "rate_limit_exceeded",
"message": "Too many requests",
"retryAfterSeconds": 10
}
}This tells clients:
request is valid
but retry laterUse One End-to-End Request Deadline
Suppose the API request budget is:
3 secondsThat does not mean every dependency gets 3 seconds.
For example:
Total request budget
3,000 ms
Database
800 ms
Inventory service
500 ms
Payment provider
1,000 ms
Response margin
300 msDependencies must fit inside the whole request budget.
Why Deadlines Matter
Without deadlines:
payment provider slow
↓
requests remain open
↓
Node connections accumulate
↓
database connections stay occupied
↓
memory grows
↓
system collapsesA timeout protects capacity.
Propagate Cancellation
Imagine the API times out after:
3 secondsbut:
database query continues for 20 seconds
payment request continuesThe user is already gone, but the system still consumes resources.
Where supported, propagate:
AbortSignal;through:
controller
application service
database adapter
HTTP clientso abandoned work can stop.
Retries Need Rules
Retries are not always good.
Retry only when:
failure may be temporary
operation is idempotent
deadline remains
retry budget allows itExamples of potentially retryable failures:
temporary network error
connection reset
503 from dependency
short timeoutExamples usually not worth retrying:
validation failure
403
payment declined
invalid API key
domain conflictAvoid Retry Multiplication
Imagine:
Mobile client
→ retries 3 times
API
→ retries payment 3 times
Payment adapter
→ retries HTTP 3 times
Queue worker
→ retries 3 timesOne original request can create:
3 × 3 × 3 × 3
= 81 attemptsduring an outage.
This is called:
retry amplificationRetries need one coordinated budget.
The retry-storm guide follows this feedback loop through timeouts, exponential backoff, jitter, and circuit breakers.
Backpressure and Load Shedding
For a deeper treatment of streams, bounded promise concurrency, API admission, and queues, see backpressure in Node.js systems.
Suppose export workers can process:
50 jobs/minutebut the API accepts:
500 jobs/minuteAn unlimited queue does not solve the problem.
It creates:
larger and larger delayEventually:
jobs are hours old
memory/storage grows
recovery takes foreverUse Bounded Queues
For example:
Queue maximum
= 1,000 jobsOnce capacity is reached:
reject new low-priority workor:
throttle producersThis prevents overload from turning into an endless backlog.
Protect Important Work First
Suppose the system is overloaded.
Prioritize:
create order
payment statusover:
recommendations
bulk exports
analyticsExample:
Keep:
order creation
payment status
Defer:
confirmation email
Reject:
optional exportThis is load shedding.
Sometimes a Fast 503 Is Better
Suppose every request takes:
30 secondsthen times out.
The system wastes:
database connections
CPU
memory
network socketsA quick:
503 Service Unavailable
Retry-After: 10may be healthier.
Failing quickly can preserve the rest of the system.
Every Request Should Be Traceable
Generate or accept a trusted request ID.
For example:
req-01K5...Return it:
X-Request-Id: req-01K5Use the same ID in logs.
Structured Logging Example
logger.info({
message: "request completed",
requestId,
traceId,
route: "POST /orders",
tenantTier: subject.tenantTier,
status: 201,
durationMs,
idempotencyOutcome: "created",
});This is easier to search than:
"Request worked."Do Not Put IDs Into Metric Labels
Bad:
http_requests_total{
userId="user-123"
}or:
{
requestId="req-..."
}These values have huge cardinality.
Metrics systems such as Prometheus work better with bounded dimensions.
Good:
route
method
status class
error class
tenant tierMeasure RED Signals
For every route template, measure:
Rate
Errors
DurationFor example:
POST /orders
requests/sec
error %
p50
p95
p99These are the classic RED metrics.
Also Measure Application-Specific Signals
Useful Order API metrics:
validation failures
authorization failures
idempotency replays
idempotency conflicts
rate-limit rejections
database pool wait
payment latency
payment timeout rate
queue depth
oldest job ageThese help explain why generic latency changed.
Distributed Traces
One request may flow:
HTTP request
↓
database
↓
inventory service
↓
payment provider
↓
queueDistributed tracing connects those operations.
During an incident you can see:
Total request = 2.8 sec
DB = 100 ms
Inventory = 200 ms
Payment = 2.3 secNow the bottleneck is clear.
Document Behavior, Not Just JSON
OpenAPI is useful for:
routes
methods
request schema
response schema
authenticationBut production contracts include semantics that schemas cannot fully describe.
Document things such as:
ordering guarantees
idempotency requirements
pagination behavior
ETag semantics
retry behavior
eventual consistency
rate-limit behavior
deprecation rulesFor example:
Idempotency-Key must remain stable
for retries of one logical order creation.A schema alone cannot communicate that correctly.
Keep Documentation Close to Implementation
Documentation that is manually maintained but never tested will drift.
Prefer:
generated OpenAPIor:
OpenAPI validated in CIso changes to routes or schemas are detected.
Evolve the API Additively
Good API evolution prefers changes like:
add optional response field
add optional request field
add new endpoint
add new operationThese are often backward-compatible.
Avoid silently doing:
number
→ stringor:
optional
→ requiredor:
total means cents
→ total means dollarsThose can break existing clients.
Version Only When Necessary
Before creating:
/v2ask:
Can old and new behavior
coexist safely?If yes:
prefer additive migrationIf no:
create an explicit new versionThen:
document migration
measure old usage
announce deprecation
set sunset policy
remove deliberatelyThe API versioning guide develops that lifecycle, including expand-and-contract changes, compatibility tests, and deprecation.
Test the Public Contract
Unit tests alone are not enough.
You need tests against actual HTTP behavior.
For POST /orders, test at least:
1. Valid creation returns 201.
2. Response contains Location.
3. Wrong content type returns 415.
4. Oversized body returns 413.
5. Malformed JSON returns 400.
6. Invalid fields return stable validation error.
7. Missing authentication returns 401.
8. Cross-tenant access cannot leak data.
9. Unknown fields cannot set server-owned values.
10. Idempotency retry creates only one order.
11. Same key + different payload is rejected.
12. Stale If-Match returns 412.
13. Invalid domain transition returns 409.
14. Pagination remains deterministic.
15. If-None-Match can return 304.
16. Rate limiting returns 429.
17. Dependency timeout respects request deadline.Test Races, Not Just Happy Paths
Production bugs often appear when two things happen at once.
Test:
two idempotent requests arrive togethertwo users update same versiontwo cancellation requests racetimeout occurs during paymentConcurrency behavior should be part of the contract.
Consumer Contract Tests
If several known clients use the API:
web
mobile
partner
internal workerconsumer-driven contract tests can protect their expectations.
But they should verify public behavior.
Do not let a consumer contract lock down:
internal class names
database tables
private implementationPutting Everything Together
A production POST /orders request may look like:
Incoming HTTP request
↓
Request ID
↓
Body-size protection
↓
Content-Type validation
↓
Authentication
↓
Authorization
↓
JSON parsing
↓
Schema validation
↓
Explicit request mapping
↓
Idempotency check
↓
PlaceOrder use case
↓
Domain validation
↓
Database transaction
↓
Payment call
↓
Persist result
↓
Structured response
↓
Metrics + logs + tracesThis is much more than:
controller
→ service
→ databasebecause production behavior includes failure and recovery.
Where Different Responsibilities Belong
A useful separation is:
HTTP layer
Content type
body limits
headers
HTTP status
request ID
schema validation
response formattingApplication layer
authorization policy
use case orchestration
idempotency workflow
business actionsDomain layer
business invariants
order transitions
pricing rulesInfrastructure
PostgreSQL
Redis
Stripe
email
queue
external APIsThe boundaries do not have to be perfect.
The goal is to stop protocol and infrastructure details from taking over business logic.
A Simple Mental Model
When designing an endpoint, ask these questions.
1. What resource or business action
does this endpoint represent?Then:
2. What input is accepted?Then:
3. Who is the caller?Then:
4. Are they allowed to do this?Then:
5. What business rules apply?Then:
6. What happens if the client retries?Then:
7. What happens if two requests race?Then:
8. What happens if a dependency is slow?Then:
9. What happens if the system is overloaded?Then:
10. How will we debug this in production?Then:
11. How can this contract evolve safely?If you can answer those questions, your endpoint is much closer to production-ready.
Production Checklist
Before shipping a REST API:
□ Model resources clearly.
□ Use HTTP methods deliberately.
□ Define safe and idempotent behavior.
□ Validate Content-Type.
□ Limit body size.
□ Parse JSON safely.
□ Validate schemas at runtime.
□ Reject or control unknown fields.
□ Prevent mass assignment.
□ Map request DTOs explicitly.
□ Derive tenant identity from auth.
□ Scope database reads by tenant.
□ Authorize the actual resource.
□ Keep domain validation in the use case.
□ Use one stable error format.
□ Use stable machine-readable error codes.
□ Avoid leaking internal failures.
□ Require idempotency for retryable POSTs.
□ Protect downstream side effects
with stable keys.
□ Use ETags for conditional reads.
□ Use If-Match for optimistic updates.
□ Use 412 for stale versions.
□ Use 409 for domain conflicts.
□ Paginate every large collection.
□ Use deterministic ordering.
□ Bound page size.
□ Define cache behavior explicitly.
□ Prevent private-response cache leaks.
□ Use 202 for long-running work.
□ Persist async operations before
acknowledging them.
□ Apply rate limits by resource.
□ Add concurrency limits for
expensive long-running work.
□ Define request deadlines.
□ Propagate cancellation.
□ Retry only safe transient failures.
□ Avoid retry amplification.
□ Bound queues.
□ Apply load shedding under saturation.
□ Generate or propagate request IDs.
□ Use structured logs.
□ Measure RED metrics.
□ Avoid high-cardinality metric labels.
□ Add distributed tracing.
□ Document API semantics.
□ Validate OpenAPI in CI.
□ Prefer backward-compatible evolution.
□ Version breaking contracts explicitly.
□ Test retries, races, failures,
and old client behavior.Conclusion
A production-ready REST API is not simply:
GET works
POST works
PATCH worksIt is an API whose behavior stays predictable when things go wrong.
A strong API answers:
What if input is invalid?
What if the caller is unauthorized?
What if the request is retried?
What if two clients update at once?
What if the database is slow?
What if payment times out?
What if the queue is full?
What if the client asks for
100,000 rows?
What if a proxy caches
private data?
What if we need to change
the contract next year?For our Order API, production readiness came from treating all of these as part of the design:
HTTP semantics
validation
authentication
authorization
domain rules
idempotency
optimistic concurrency
pagination
caching
asynchronous operations
rate limiting
deadlines
retries
backpressure
load shedding
observability
contract evolutionThe main lesson is:
Production-ready APIs are designed around failure, retries, concurrency, and change—not only the happy path.
If clients can retry safely, concurrent requests cannot silently overwrite one another, tenant data stays isolated, expensive work is bounded, failures are observable, and the contract can evolve without surprising consumers, then the API is much closer to being truly production-ready.
References
Related
Written by
Faisal
Software engineer writing about backend systems, Node.js, system design, scalable applications, and modern web and mobile development.