v1 enforces a rate limit of 300 requests per minute, counted per calling IP address across every v1 endpoint. Requests over the ceiling are rejected with 429 Too Many Requests and a Retry-After header.
This limit reached production on 1 September 2026. If you built against the earlier documentation, which stated that v1 had no ceiling, your integration has no 429 handling — add it.

How the limit works

Read Retry-After

A 429 carries Retry-After with the whole number of seconds until the window resets. Wait that long before retrying.
v1 sends no X-RateLimit-Limit, X-RateLimit-Remaining or X-RateLimit-Reset headers. You cannot read your remaining allowance from a successful response — track your own send rate. Those headers arrive with v2.

The window is fixed, not rolling

The allowance resets on a wall-clock minute boundary rather than sliding. Two bursts either side of a boundary both succeed, so a short span can carry close to 600 requests without a 429. Don’t build on that — spread your traffic instead of aiming at the boundary.

The counter is per IP, not per key

The partition is the calling address, not your API key. Two consequences:

Shared egress shares the bucket

Every process behind one NAT or egress IP draws on the same 300. A batch job and your live integration will throttle each other.

Separate egress raises your effective ceiling

Callers on distinct addresses each get their own allowance. This is a property of the implementation, not a supported way to buy headroom.
The counters are held in each API instance rather than shared between them, so the ceiling you actually observe is 300 multiplied by however many instances are serving traffic — a number that changes with scaling and is not published. Calibrating your client to the throughput you measure will break the next time production scales in. Build to 300 per minute.

Building considerately

Page, don't poll hard

Search endpoints take start and take. Pull a page at a time rather than looping at speed over the whole marketplace.

Back off on failure

On 429, honour Retry-After. On 5xx or a timeout, retry with exponential backoff — but read the retry guidance first, because v1 has no idempotency keys.

Cache what's stable

Group membership and NDC statistics change slowly. Re-reading them on every operation is wasted traffic.

Batch NPI updates

groups/npi-update takes an array. Send one call with many NPIs, not many calls.

What changes in v2

v2 replaces the per-IP ceiling with per-credential and per-pharmacy limits, with standard headers on every response — X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset — and Retry-After on a 429. Search requests count double against the allowance, because they’re the expensive ones.
Exact v2 ceilings will be published with the v2 documentation at general availability.