— The AI Factory

A real AI Factory.
On South African soil.

Plenty of companies will sell you "AI." Far fewer own the infrastructure it runs on. The FirstCoreAI AI Factory is physical compute we operate ourselves — GPU nodes running production inference, evaluated against the best frontier APIs so we know exactly where each model earns its place.

— What's inside

Compute, models, endpoints — all under one roof.

COMPUTE

Current-generation Blackwell-class GPUs

A dedicated GPU cluster built on current-generation NVIDIA Blackwell-class hardware, paired with AMD EPYC compute and over a terabyte of system RAM per node. Distributed serving tuned for the throughput and concurrency of production traffic rather than leaderboard runs.

MODELS

Six production endpoints, kept current

An 80B mixture-of-experts coding model (131K context), a 31B general chat model (128K context), vision-language, speech-to-text, multilingual embeddings, and a retrieval reranker — all live on the Factory today. Where a workload genuinely needs a frontier API, we route to one.

ENDPOINTS

Private, OpenAI-compatible APIs

Every model is delivered as an OpenAI-compatible HTTP API — chat completions, embeddings, audio transcriptions, rerank — bearer-token authenticated, TLS-terminated at the cluster edge. Any SDK that talks to OpenAI talks to the Factory without code changes.

STORAGE & FABRIC

Multi-TB NVMe, high-throughput in-band fabric

A multi-terabyte NVMe array holds the shared model cache, NFS-mounted across the cluster — sized for production model swapping without re-downloads. A high-throughput in-band fabric between nodes carries the distributed-serving traffic; out-of-band management is segregated.

— Performance proof

Stress-tested. Numbers from our own rig.

Measured on the production cluster under real serving conditions — not lifted from a vendor slide.

~0.6s
Time to first token · chat & coding

Median from request to first streamed token, staying under a second at the median during concurrent bursts.

271
Concurrent requests · peak tested

Sustained across the cluster at the peak of the full stress run.

0
Errors · across all stress tests

Every request in the stress-test suite completed cleanly, at every concurrency level tested.

— The stack

Built on the tools that scale AI for real.

Kubernetes for orchestration. Run.ai for GPU-aware scheduling and project-level resource quotas. vLLM for the inference runtime, exposing an OpenAI-compatible API on every model. HAProxy and a Knative-class gateway terminate TLS at the cluster edge; pod-level OAuth2 proxies validate each request.

Bearer-token authenticated. No client SDK changes when migrating from OpenAI. Model storage on a shared NVMe NFS cache fed from HuggingFace Hub. Grafana and Prometheus for the operations side. The same engineers who run the Factory are the people you talk to when something needs tuning.

— How we choose

The credibility detail.

We don't guess. Every model we put into production is benchmarked on our own hardware against frontier APIs for the specific job — accuracy, latency, throughput under real concurrency, and cost per unit of work. Sometimes a self-hosted open model wins outright. Sometimes a frontier API is worth the premium. Either way, you see the evidence before you commit.

— Why it matters

What you actually get out of it.

  • Performance you can plan around — dedicated capacity rather than a shared queue.
  • Costs you can forecast — rand-denominated and workload-based, so a prompt tweak doesn't move your bill.
  • Data that never leaves — every inference happens inside the perimeter.
  • A team that built it — when something needs tuning, the people who run the Factory are the people you talk to.

Want to see the Factory in action?

We'll walk you through the hardware, the model catalogue, and a workload running live.