The compute layer.
Sovereign by default.
InfraAI is the AI Factory delivered as a service. Private inference endpoints, dedicated GPU capacity, and a curated catalogue of open-weight models — all hosted in South Africa, all consumed through clean, OpenAI-compatible APIs.
The InfraAI stack, end to end.
- Private inference endpoints — OpenAI-compatible HTTP APIs, hosted locally. Model in, answer out. Nothing leaves the country.
- Six production endpoints, live today — coding (80B MoE), general chat (31B), vision-language, speech-to-text, multilingual embeddings, and a retrieval reranker.
- Dedicated and shared GPU options — scale capacity to the workload, not the marketing.
- Frontier routing when needed — one integration, the right model per call. We tell you when a request leaves the country.
- Rand-denominated commercials — predictable, forecastable, no FX surprise.
Drop-in OpenAI compatibility.
Authenticate with OAuth2 client credentials, exchange for a short-lived bearer token, and call any of the standard OpenAI endpoints —
/v1/chat/completions, /v1/embeddings, /v1/audio/transcriptions, /v1/rerank.
Any SDK that works with OpenAI works with InfraAI without code changes. The token endpoint and your specific endpoint URLs are issued at onboarding.
Where InfraAI is the right call.
High-volume document and call processing, retrieval and RAG backends, internal assistants, classification and extraction — and any workload where data sensitivity or cost makes a foreign API a bad fit.
Got a workload that shouldn't be leaving the country?
We'll spin it up on the Factory and show you what sovereign feels like.
