now building · Reference Platform
drag to orbit · click to switch build
Private AI · runs on your own computers · works offline

Use powerful AI
without giving away your data.

Torii runs AI on the computers in your own building — not in someone else's cloud. Your files never leave. It even works with the internet unplugged. And the decisions that matter — who gets paid, how much, who qualifies — are made by rules you approve, never by an AI that might guess wrong.

100%runs on hardware you own
0files sent to the cloud
0guesses on money decisions
works with no internet
Stage 1 · ProblemWhy AI is hard to trust today.
The problem

Using AI today means giving something up

Every company wants AI. But right now, using it forces three bad trade-offs.

01

Your data goes to strangers

To use most AI, you upload your private files to a giant company's computers — often in another country. For a hospital, a bank, or a government office, that alone is a dealbreaker.

02

The AI guesses — and is sometimes wrong

Today's AI makes things up. You can't let a machine that guesses decide who gets a benefit, a loan, or a medical payment — and when it's wrong, nobody can explain why.

03

You get locked in

One vendor, one chip brand, one cloud. Prices climb, you can't see what it's really doing, and you can't leave without starting over.

Torii removes all three

Run the AI on your own machines. Keep your data in the building. Let rules you approve make the real decisions — and prove every one of them.

Stage 2 · SolutionTorii — private AI you own. Try it, see what it needs, see how it decides.
In one sentence

A complete AI system that runs entirely on hardware you own — where a rulebook you approve makes the decisions, not the AI.

The AI reads documents, drafts replies, and explains things. But the moment a real decision is made — an amount, an approval, an eligibility — a rulebook your own people wrote and signed decides it. Nothing is sent outside. Nothing is guessed. Everything is written down and can be checked later.

Live · this part really runs on your device

Try it right here — no install, no internet

The large-scale infrastructure below is drawn as a 3D model. This part — and the live monitor just below — is real. Modern browsers can run small AI models directly on your computer's own graphics chip (WebGPU). Nothing is uploaded — the model runs on your machine, and it keeps working with the internet unplugged. Press the button and we'll check, right now, what your computer can do.

reads your device only · nothing leaves the page
▶ The live app really runs here Pick a field (cooking, coding, construction…), and a small AI model downloads to your browser and runs on your own GPU — grounded in a knowledge base you build, and it keeps working with the internet unplugged.
Want it deeper in your browser? The Torii · Local AI Analyzer extension reads your exact hardware and runs bigger on-device models — currently in Chrome Web Store review.
The journey · add power as you go

Start free in your browser — grow into a full sovereign platform

Level 1
In your browsernow · offline · free

A small on-device model. Summarize, extract, draft, classify — even with the internet unplugged.

Level 2
Your computer's GPUmeasured → suggested

Use your machine's full graphics card to run a bigger model — adds translation, longer documents, stronger reasoning.

Level 3
Your serverson-prem

The full platform on your own servers — 70B models, agents, and the rulebook that makes decisions.

Level 4
Rented frontierrent-to-own

The very biggest open models (Kimi K2, GLM-5.2) on rented 8× H200 you later buy and bring on-site.

Live monitor · queried from your browser

What's actually running — right now

Not a simulation. The components below are live: on-device AI runs in your browser via WebGPU, and the decision core, gateway and knowledge base run on the edge. The numbers were just fetched from this page.

Edge core
checking…
Rule packs
human-attested · live
Live decision
decided by the rulebook
Knowledge base
OKF · Vectorize + R2
On-device modelsWebGPU · runs in your browser
The live app offers 1B & 3B. The Local AI Analyzer extension unlocks bigger models to match your hardware — an 8B ran on an Apple M4 Pro.
Knowledge base accessOKF · live-pinged
Build & curate domains in the live app ↗. Endpoints: /okf · /knowledge/pack · /knowledge/chat · /knowledge/agent.
How decisions are made

The AI helps.
But it never touches the money.

Think of a bank teller and a rulebook. The AI is the teller — it reads your documents, fills in the forms, and explains what's going on. But the rulebook, written and signed off by your own staff, decides the actual amount or approval. The teller can never change the rulebook. If anything is unclear, the AI stops and hands it to a person. And every decision gets a tamper-proof receipt you can check later — change a rule and the system refuses it.

Approve ¥30,000 · decided by the rulebook · not the AI
What you need to run it

It runs on ordinary hardware you already own

No magic boxes and no special cloud account. If you have servers and a few graphics cards, you can run it.

Servers
Standard data-centre servers — the same kind you already run. Nothing exotic.
GPUs
Graphics cards of any brand. It shares each one across many jobs, so you need far fewer.
On-site
Your own building or data centre — or even a single machine with no internet at all.
No cloud
Nothing depends on Amazon, Google, or anyone else's computers. It's all yours.
Match the model to the machine

Bigger models need more hardware — here's exactly how much

An image model like Stable Diffusion runs on one graphics card. A frontier text model like Kimi K2 or GLM-5.2 needs eight of the most powerful cards (H200) in a single machine. This chart is honest about what each one takes.

Imagediffusion
Stable Diffusion XL, SD3, FLUX.1 — generates pictures in seconds
Own it · 1 GPU (~$10k)
Edge≤ 8B params
Qwen 7B, Swallow 8B — chat, extraction
Own it · 1 GPU (~$10k)
Standard13–34B
Qwen 32B — analysis, drafting
Own it · 1–2 GPUs (~$25k)
Large70B
Llama 70B, Swallow 70B — deep reasoning
Own it · 2–4× H100 (~$120k)
Very large~671B · MoE
DeepSeek-V3 / R1 — frontier reasoning
Own or rent-to-own · 8× H200
Frontier~1T · MoE
Kimi K2, GLM-5.2 — the biggest open models
× nodes
Rent-to-own → move on-prem
What it does for you

Four things it gives you — in plain terms

The small grey line under each is the open-source software that does the work.

01

Get the most out of your GPUs

HAMi · K8s DRA · NVIDIA KAI

Graphics cards are expensive and usually sit half-idle. This shares each card across many jobs and pushes it to as much as 90% use — so you buy far fewer of them, from any brand.

02

Run the right model for your hardware

vLLM · SGLang · LoRA · MLflow

A small model runs on one graphics card; the biggest open models need a rack of them. Torii runs any of them — and tells you up front exactly what each one needs (see the chart below). Japanese-native models included; switch any time.

03

Locked down, and provable

Keycloak · OpenFGA · OpenBao · Presidio

Proper logins, permissions, and privacy filtering — plus a tamper-proof record of every action, so you can always show exactly what happened and why.

04

Pay for results, with a hard limit

gateway · OpenMeter · budget gate

You're billed for what the AI actually delivers, per use. A built-in budget cap stops spending before it ever goes over — the request is simply refused, not overspent.

See it for yourself

Spin the whole system around

An interactive 3D model of the platform — the servers, the AI, and the connections between them. You can rotate it, and switch between ten ready-made setups (from GPU sharing to a fully offline install). Type a plain request and watch the model rebuild itself.

Open the simulator ↗
The building blocks

Every part is open-source — nothing is hidden

The platform is made of these pieces — each shown in the 3D model with the job it does. decide-core and gateway are live on the edge right now; the rest are on the roadmap (marked below). Every little mark is generated from the piece's name, so no two look alike.

Us vs. the usual way

The same power — without the strings attached

The big proprietary platforms lock you to their country, their code, and their cloud. Here's the side-by-side.

 
The usual way (big cloud AI)
Torii
Sovereignty
Tied to one jurisdiction (信创)
Yours — on-prem, air-gappable, portable
Source
Closed, vendor-controlled
Open — inspect, fork, self-host
Decision path
Model in the loop
Deterministic core decides — model never on the money path
Audit
Vendor attestation
Hash-chained, PII-free, replayable by you
Lock-in
Chips, models, cloud
None — multi-vendor, any model, any substrate
How the process works

A simulation of what the system does, step by step

These are the guarantees the live decision core enforces on every call — proven and replayable, not a mock-up. Try each one yourself in the building-blocks above.

Same
Same input, same result — the decision engine follows a fixed rulebook, so it can't drift or make things up.
Refused
If a rule is tampered with, the system won't use it — the change is blocked.
Stopped
When a customer hits their spending limit, the request is turned away before it costs anything.
0
Money decisions made by the AI — the rulebook makes them, and each one is recorded.
Built on trusted open tools

Made from proven open-source parts — owned by no single company

KubernetesvLLMSGLangHAMiKeycloakOpenFGAOpenBaoTemporalMilvusApache IcebergPresidioOpenTelemetryCephZarfRustGo
Stage 3 · TokenomicsHow you pay — and how you end up owning it.
The money model

Pay for what you use. Cap what you spend. Own the hardware.

Three simple rules.

Rule 1

Pay for results

AI is billed by the token — the unit of work it does. You pay for what it delivers. Nothing idle.

per token
Rule 2

Never overspend

Every team gets a budget. Hit the cap and the next request is refused — before it costs a cent.

hard cap
Rule 3

Own the hardware

Rent the big machines, run today, then buy them and move them in-house. Rent counts toward the purchase.

rent → own
Rule 3, in detail · rent-to-own

Don't own eight H200s yet? Start renting — end up owning.

Rent a dedicated machine — yours alone, with proven high-bandwidth connectivity, not shared public cloud. Run the big model today, then buy that same hardware and move it into your building. Sovereignty on an installment plan.

STEP 1
Rent

A dedicated 8× H200 node with proven high-bandwidth connectivity — reserved for you alone.

≈ $25k / mo$0 upfront · cancel anytime
STEP 2
Run today

Serve the big model now, while your own data centre is still being fitted out.

$0 capexno capital outlay · pay monthly
STEP 3
Buy

Rent already paid is credited toward the purchase — the hardware becomes yours.

≈ $350k to ownup to 100% of rent credited
STEP 4
Move on-prem

Ship the machine into your building. Now it's fully yours, on-site.

≈ $3k / mopower & upkeep only · no more rent
Indicative example figures for an 8× H200 node — real pricing depends on provider, GPU generation, term, and configuration. The point: start with zero capital outlay, end owning the hardware.
Try it

Your AI. Your data. Your building.

Spin the system around, switch between setups, and watch the rulebook make a decision — all open-source, from top to bottom.