NVIDIA and domestic silicon pooled together under one inventory and one scheduler. Exactly one thing is exposed to the outside: an OpenAI-compatible token endpoint. Which card served the call is our problem, not yours.
No bare metal, no foundation-model training, no physical data-centre operations, no vertical applications. A clear boundary — the rest belongs to our partners.
Language, multimodal, embedding, reranking and speech, all behind an OpenAI-compatible protocol. Swap one base_url and you are connected.
Dedicated compute and scheduling domain on a long-term contract, with nothing else competing for it. Suited to steady, latency-sensitive production traffic.
Tuning at two levels, card and model: operators, parallelism, batching and KV-cache management — squeezing throughput out of the same silicon.
Weights and data stay inside the in-country node the customer nominates, with the full call path auditable — enough to satisfy data-residency requirements.
Every step on this path leaves a record in the platform. That is how we hold ourselves accountable to the partners who fund the hardware.
A compute centre is onboarded, cards join the pool by silicon family, and serial number, firmware and owner are registered.
Models land on silicon that can run them per the compatibility matrix, with scheduling domain and replica count set.
Each request is routed by proximity, idle capacity, domestic-silicon preference and compliance isolation.
Streamed back through the OpenAI-compatible endpoint, with time-to-first-token and throughput observable end to end.
Metered by token and written to the ledger; quota, rate limits and usage are live in the console.
Open-weight flagships, multimodal, embedding and reranking, speech — and we keep pace with new releases.
Heterogeneous does not mean throwing cards in a heap. Each silicon family runs the models it is actually good at, and the platform hides the difference from the caller.
Delivery is handled by ecosystem partners; the platform supplies models, compute and inference tuning.
From embodied AI and scientific computing to autonomous driving and public-sector work, our partners both consume the compute and deliver it into production. Five leading open-weight models run on our own nodes.
Fully OpenAI-compatible. Existing code keeps its structure — one line of configuration moves it across. Auth, rate limiting, metering and streaming all work out of the box.
A default key is generated on registration; create or disable more in the console at any time.
Choose from the catalog, or just call by model name — the platform routes it to suitable silicon.
curl, the OpenAI SDK and LangChain all work as-is, streaming included.
from openai import OpenAI client = OpenAI( api_key="sk-yr-YOUR-KEY", base_url="https://api.yrinvestment.net/v1", # change only this line ) resp = client.chat.completions.create( model="DeepSeek-V4.1-Flash", messages=[{"role": "user", "content": "Hello"}], stream=True, ) for chunk in resp: print(chunk.choices[0].delta.content or "", end="")
You get an API key on registration, make real calls and watch the usage post to your account.