New DeepSeek-V4.1-Flash now live · 1M context

Heterogeneous cards,
one endpoint.

NVIDIA and domestic silicon pooled together under one inventory and one scheduler. Exactly one thing is exposed to the outside: an OpenAI-compatible token endpoint. Which card served the call is our problem, not yours.

—
Cards under management
—
Daily token output
—
Models live
—
Compute centres
What the platform does

We do exactly four things

No bare metal, no foundation-model training, no physical data-centre operations, no vertical applications. A clear boundary — the rest belongs to our partners.

Model API

Language, multimodal, embedding, reranking and speech, all behind an OpenAI-compatible protocol. Swap one base_url and you are connected.

  • One key, every model
  • Streaming and function calling
  • Metered by token, quota under your control

Reserved capacity

Dedicated compute and scheduling domain on a long-term contract, with nothing else competing for it. Suited to steady, latency-sensitive production traffic.

  • Dedicated scheduling domain, isolated resources
  • Enterprise availability commitment
  • Long-term capacity with elastic overflow

Inference tuning

Tuning at two levels, card and model: operators, parallelism, batching and KV-cache management — squeezing throughput out of the same silicon.

  • Operator support across heterogeneous silicon
  • Parallelism and replica orchestration
  • End-to-end latency breakdown and attribution

Private deployment

Weights and data stay inside the in-country node the customer nominates, with the full call path auditable — enough to satisfy data-residency requirements.

  • In-country nodes with residency labelling
  • A compliant pool built on domestic silicon
  • End-to-end audit of models and calls
From card to token

How a card becomes an API call

Every step on this path leaves a record in the platform. That is how we hold ourselves accountable to the partners who fund the hardware.

01

Inventory

A compute centre is onboarded, cards join the pool by silicon family, and serial number, firmware and owner are registered.

02

Scheduling

Models land on silicon that can run them per the compatibility matrix, with scheduling domain and replica count set.

03

Routing

Each request is routed by proximity, idle capacity, domestic-silicon preference and compliance isolation.

04

Inference

Streamed back through the OpenAI-compatible endpoint, with time-to-first-token and throughput observable end to end.

05

Metering

Metered by token and written to the ledger; quota, rate limits and usage are live in the console.

Model serving

One key, every model

Open-weight flagships, multimodal, embedding and reranking, speech — and we keep pace with new releases.

See all models →
Compute network

Four centres, four silicon families, one pool

Heterogeneous does not mean throwing cards in a heap. Each silicon family runs the models it is actually good at, and the platform hides the difference from the caller.

Why we do not quote an “equivalent compute” number. Reducing different brands of silicon to one yardstick does not hold up for inference — the conversion factor moves with the model; it is not a property of the card. The same Ascend card gives perfectly usable throughput on Qwen3 and will not start at all on Llama 4, which the vendor explicitly does not support. So we list real card counts per silicon family and pair them with a “which models actually run” matrix, rather than a synthetic number that falls apart under the first question. View the compatibility matrix →
Industry solutions

Five workloads where compute demand is least negotiable

Delivery is handled by ecosystem partners; the platform supplies models, compute and inference tuning.

Ecosystem

Twenty portfolio companies,
five domains, one network

From embodied AI and scientific computing to autonomous driving and public-sector work, our partners both consume the compute and deliver it into production. Five leading open-weight models run on our own nodes.

20
Portfolio companies
12
Listed
5
Domains
40+
Customer countries
Explore the ecosystem →
Quick start

Change one base_url
and you are connected

Fully OpenAI-compatible. Existing code keeps its structure — one line of configuration moves it across. Auth, rate limiting, metering and streaming all work out of the box.

01

Register and create an API key

A default key is generated on registration; create or disable more in the console at any time.

02

Pick a model

Choose from the catalog, or just call by model name — the platform routes it to suitable silicon.

03

Make the call

curl, the OpenAI SDK and LangChain all work as-is, streaming included.

quickstart.py
from openai import OpenAI

client = OpenAI(
    api_key="sk-yr-YOUR-KEY",
    base_url="https://api.yrinvestment.net/v1",  # change only this line
)

resp = client.chat.completions.create(
    model="DeepSeek-V4.1-Flash",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)

for chunk in resp:
    print(chunk.choices[0].delta.content or "", end="")
What a single call passes through. API key check → quota check → routing policy match → model serving instance → streamed response → usage written to the ledger. Every step is queryable in the console, and the response itself carries the silicon family and compute centre that served it.

Try it now

You get an API key on registration, make real calls and watch the usage post to your account.

Get started → Read the docs