Resources · Blog
Notes on sovereign inference.
What we are learning building an NPU-native gateway in Pakistan, and what it means for the teams that call it. Short, practical, and free of benchmarks we have not run.
Page 1 of 2
8 min read
What an LLM gateway does, and what changes when it runs in Pakistan
An LLM gateway is one API in front of many models, holding keys, spend caps, routing and cost in one place. What changes when that gateway runs inside Pakistan.
Read the post9 min read
An AI compliance checklist for banks in Pakistan
The questions a bank's risk and compliance function asks before an AI feature ships in Pakistan: where data is processed, what is retained, and who can prove it.
Read the post7 min read
What a million tokens actually costs in rupees
A million tokens has two prices, not one, and output usually dominates the bill. How to turn a real workload into a rupee figure before you write any code.
Read the post7 min read
Prompt logging: what your AI provider keeps after the answer
Most inference providers retain prompts and completions for a period. What gets kept, how to find out before your data is in it, and what P/9 records instead.
Read the post7 min read
Rate limits, spend caps and expiry, one key at a time
A key that cannot outspend its budget, outlive its purpose or reach the wrong model is a smaller incident than one that can. The four limits on a P/9 API key.
Read the post9 min read
What is a sovereign AI gateway?
A sovereign AI gateway is one API in front of many models where the operator can tell you which country served each request, in which currency you were billed, and what was retained. Here is what that means in practice.
Read the post7 min read
Paying for an AI API from Pakistan is the hard part
The model call is easy. Getting a foreign card to work, holding a dollar balance and explaining the exchange-rate line to finance is what stalls projects. What changes when the invoice is in rupees.
Read the post8 min read
Pakistan's data localisation rules, and what they mean for AI
The Personal Data Protection Bill expects critical personal data to be processed inside Pakistan. Sending it to an inference endpoint abroad is a processing decision, and it is the one most AI projects never make explicitly.
Read the post7 min read
Choosing an inference provider when location is a requirement
The aggregators route to whichever capacity is cheapest, and that is the right answer for most teams. It stops being the right answer when where a request ran is a requirement rather than a preference.
Read the post