Torii runs AI on the computers in your own building — not in someone else's cloud. Your files never leave. It even works with the internet unplugged. And the decisions that matter — who gets paid, how much, who qualifies — are made by rules you approve, never by an AI that might guess wrong.
Every company wants AI. But right now, using it forces three bad trade-offs.
To use most AI, you upload your private files to a giant company's computers — often in another country. For a hospital, a bank, or a government office, that alone is a dealbreaker.
Today's AI makes things up. You can't let a machine that guesses decide who gets a benefit, a loan, or a medical payment — and when it's wrong, nobody can explain why.
One vendor, one chip brand, one cloud. Prices climb, you can't see what it's really doing, and you can't leave without starting over.
Run the AI on your own machines. Keep your data in the building. Let rules you approve make the real decisions — and prove every one of them.
The AI reads documents, drafts replies, and explains things. But the moment a real decision is made — an amount, an approval, an eligibility — a rulebook your own people wrote and signed decides it. Nothing is sent outside. Nothing is guessed. Everything is written down and can be checked later.
The large-scale infrastructure below is drawn as a 3D model. This part — and the live monitor just below — is real. Modern browsers can run small AI models directly on your computer's own graphics chip (WebGPU). Nothing is uploaded — the model runs on your machine, and it keeps working with the internet unplugged. Press the button and we'll check, right now, what your computer can do.
A small on-device model. Summarize, extract, draft, classify — even with the internet unplugged.
Use your machine's full graphics card to run a bigger model — adds translation, longer documents, stronger reasoning.
The full platform on your own servers — 70B models, agents, and the rulebook that makes decisions.
The very biggest open models (Kimi K2, GLM-5.2) on rented 8× H200 you later buy and bring on-site.
Not a simulation. The components below are live: on-device AI runs in your browser via WebGPU, and the decision core, gateway and knowledge base run on the edge. The numbers were just fetched from this page.
Think of a bank teller and a rulebook. The AI is the teller — it reads your documents, fills in the forms, and explains what's going on. But the rulebook, written and signed off by your own staff, decides the actual amount or approval. The teller can never change the rulebook. If anything is unclear, the AI stops and hands it to a person. And every decision gets a tamper-proof receipt you can check later — change a rule and the system refuses it.
Approve ¥30,000 · decided by the rulebook · not the AINo magic boxes and no special cloud account. If you have servers and a few graphics cards, you can run it.
An image model like Stable Diffusion runs on one graphics card. A frontier text model like Kimi K2 or GLM-5.2 needs eight of the most powerful cards (H200) in a single machine. This chart is honest about what each one takes.
The small grey line under each is the open-source software that does the work.
Graphics cards are expensive and usually sit half-idle. This shares each card across many jobs and pushes it to as much as 90% use — so you buy far fewer of them, from any brand.
A small model runs on one graphics card; the biggest open models need a rack of them. Torii runs any of them — and tells you up front exactly what each one needs (see the chart below). Japanese-native models included; switch any time.
Proper logins, permissions, and privacy filtering — plus a tamper-proof record of every action, so you can always show exactly what happened and why.
You're billed for what the AI actually delivers, per use. A built-in budget cap stops spending before it ever goes over — the request is simply refused, not overspent.
An interactive 3D model of the platform — the servers, the AI, and the connections between them. You can rotate it, and switch between ten ready-made setups (from GPU sharing to a fully offline install). Type a plain request and watch the model rebuild itself.
Open the simulator ↗The platform is made of these pieces — each shown in the 3D model with the job it does. decide-core and gateway are live on the edge right now; the rest are on the roadmap (marked below). Every little mark is generated from the piece's name, so no two look alike.
The big proprietary platforms lock you to their country, their code, and their cloud. Here's the side-by-side.
These are the guarantees the live decision core enforces on every call — proven and replayable, not a mock-up. Try each one yourself in the building-blocks above.
Three simple rules.
AI is billed by the token — the unit of work it does. You pay for what it delivers. Nothing idle.
Every team gets a budget. Hit the cap and the next request is refused — before it costs a cent.
Rent the big machines, run today, then buy them and move them in-house. Rent counts toward the purchase.
Rent a dedicated machine — yours alone, with proven high-bandwidth connectivity, not shared public cloud. Run the big model today, then buy that same hardware and move it into your building. Sovereignty on an installment plan.
A dedicated 8× H200 node with proven high-bandwidth connectivity — reserved for you alone.
Serve the big model now, while your own data centre is still being fitted out.
Rent already paid is credited toward the purchase — the hardware becomes yours.
Ship the machine into your building. Now it's fully yours, on-site.
Spin the system around, switch between setups, and watch the rulebook make a decision — all open-source, from top to bottom.