Approved AI Processors
1. How to read this page
AI inference can involve both a broker and a serving processor. The broker routes a request. The serving processor runs the model endpoint. PRAMAAN evaluates both layers before allowing a route for customer matter content.
This public page summarizes the approved AI inference posture. The in-product registry and per-run audit trail are the operational source for the exact route used on a given inference.
2. Broker / routing layer
OpenRouter. PRAMAAN uses OpenRouter as the broker/routing layer for current approved global AI inference routes. PRAMAAN configures brokered requests with no-training controls and provider allowlists for approved routes.
When a route uses OpenRouter, selected prompts, instructions, source excerpts, metadata, and generated outputs are transmitted to OpenRouter for routing and then to the selected serving processor.
3. Serving processors currently approved for global inference
- Amazon Bedrock - Claude model family; brokered route through OpenRouter; training excluded; global region class.
- Google Vertex AI - Gemini model family; brokered route through OpenRouter; training excluded; global region class.
- Azure OpenAI - GPT model family; brokered route through OpenRouter; training excluded; global region class.
These processors may change as model quality, availability, contractual terms, and residency options change. Material changes should be reflected in PRAMAAN's approved-processor list and audit records.
4. OpenRouter-accessible provider network
OpenRouter publishes a broader provider network. The providers below were visible on OpenRouter with no-training markers when reviewed. They are not all enabled for every firm or every matter by default; PRAMAAN uses approved allowlists, firm policy, and per-run audit records to identify the exact route used.
- Amazon Bedrock (US) - No training
- Google Vertex AI (US) - No training
- Azure OpenAI (US) - No training
- NovitaAI (US) - No training
- DeepInfra (US) - No training
- AtlasCloud (US) - No training
- Parasail (US) - No training
- Z.ai (SG) - No training
- Weights & Biases (US) - No training
- Fireworks (US) - No training
- Groq (US) - No training
- Morph (US) - No training
- Together (US) - No training
- AkashML - No training
- Inceptron (SE) - No training
- ModelRun (US) - No training
- DigitalOcean - No training
- DekaLLM (ID) - No training
- Venice (US) - No training
- Nebius Token Factory (NL) - No training
- Moonshot AI (SG) - No training
- Decart (US) - No training
- NextBit (ES) - No training
- Cerebras (US) - No training
- Phala (US) - No training
- Wafer (US) - No training
- io.net (US) - No training
- SambaNova (US) - No training
- Perceptron (US) - No training
- Perplexity (US) - No training
- Baseten (US) - No training
- MARA (US) - No training
- Seed (SG) - No training
- Mancer - No training
- Ionstream (US) - No training
- Inception - No training
- Clarifai (US) - No training
- Infermatic - No training
- Reka AI - No training
- Relace - No training
Provider policy can vary by endpoint, model family, route, and contractual setting. If a route is materially added, removed, or changed for customer matter content, PRAMAAN should update the approved-processor record and preserve the exact provider and region in the audit trail.
5. Training posture
PRAMAAN-approved AI inference routes are configured so customer matter content is not used to train third-party foundation models.
6. Firm-controlled inference policy
Firms with India-only inference requirements can request an India-resident inference policy from the settings panel. The firm chooses its inference policy, can change it, and PRAMAAN records policy changes plus per-run provider and region details in the audit log.
Model quality and availability can differ by residency policy. Global routes may use stronger model families than routes constrained to a single country.
7. Prompt caching
To answer faster and at lower cost, AI providers reuse work they have already done on repeated parts of a request. When two requests begin with the same context — the same instructions, the same source excerpt — the provider can hold the processed form of that shared opening in a temporary store called a cache, and skip recomputing it next time. This is prompt caching.
What is cached is the repeated opening of a request, held as provider-side cache state, not as a readable copy of your file. It is used only to speed up the provider's own responses. It is stored encrypted, expires automatically, and is not used to train models.
Each cache entry has a time limit, often called a TTL (time to live). Depending on the provider and settings, this runs from about five minutes up to twenty-four hours or more. The limit is rolling: on most providers, each time the cached context is reused, the timer resets. So under steady use, the same context can stay in a provider's cache for the length of a working session, and on some providers it cannot be cleared manually before it expires on its own.
The caching behaviour of the current approved routes, with links to each provider's official documentation:
- OpenRouter (broker). Passes caching through to the serving provider and keeps related requests on the same provider endpoint so the cache can be reused; it does not run a separate cache of its own. Documentation.
- OpenAI. Standard cache stays active for about 5–10 minutes of inactivity, up to a maximum of one hour; optional extended caching keeps context for up to 24 hours. Documentation.
- Anthropic (Claude). 5-minute default, with an optional 1-hour setting; each reuse of the cached content refreshes the timer. Documentation.
- Google (Gemini / Vertex AI). Automatic caching lasts a few minutes; a manually created context cache defaults to 60 minutes and can be set to a different length. Documentation.
- AWS Bedrock. Each successful cache hit resets the TTL; many models use a 5-minute TTL, with a 1-hour option on some. Documentation.
- Azure OpenAI. Standard cache is cleared within 5–10 minutes of inactivity and within one hour of last use; extended caching keeps context for up to 24 hours and is the default on newer models. Documentation.
- xAI (Grok). Caching is automatic; xAI does not publish a fixed retention time, and cache entries are evicted based on system load. Documentation.
Caching terms can change and can vary by model and route. If a route materially changes how customer matter content is cached, PRAMAAN should update this disclosure and preserve the exact provider and region in the audit trail.
8. Prompt caching, cross-border processing, and DPDP
India's Digital Personal Data Protection Act, 2023, and the Digital Personal Data Protection Rules, 2025 (notified on 13 November 2025, with core compliance obligations phasing in through 2027) govern how personal data is processed. As part of this disclosure, PRAMAAN records that an approved processor may hold matter context transiently in a prompt cache while a request is being processed, as described in section 7.
This caching runs on the same approved processors that perform the inference. Like the inference itself, it may take place outside India, including in the United States or other global regions, consistent with the AI inference described on this page and in the AI Processing Disclosure. This page discloses how the service works; it is not legal advice.
9. Questions
For processor questions or diligence requests, contact legal@pramaan.io.