Your run died at the cap.
Your other accounts sat idle.
Nubivola plans a job once, then schedules it across all the AI capacity you already pay for — every API key, every cloud account, every batch lane. When one ceiling hits, the work moves to another. When every window is shut, the run waits instead of dying.
Your keys never leave your account. Every key is encrypted in a vault, serves only your runs, and is never pooled, resold, or marked up — there is no share switch anywhere in the product, and no code path that could hand your key to anyone else's work.
Plan with a model. Schedule with code.
One frontier call turns your goal into a typed graph of subtasks. After that no model is in the loop — the scheduler is deterministic, so coordination costs zero tokens and never routes your work to the wrong worker.
Your own keys, in class-and-function order, with batch and off-peak lanes preferred for work that can wait. A 429 cools the key pool-wide and the call retries on the next best key — possibly another vendor, because the plan asks for a model class, not a model name. No quota anywhere → the run pauses and resumes when your windows roll.
Contained by construction
Every key serves only your runs. There is no sharing feature to misconfigure and no path that attaches your key to another account's work — it's an invariant tested in CI, not a policy we promise.
Right model, right job
Eight canonical functions — architecture, complex coding, bulk coding, research, review, docs, diagrams, data — each routed to a key that's good and cheap at that. Frontier attention only where the difficulty lives.
Checked before you see it
Deterministic gates, then an independent model that didn't write the work, then a bounded repair loop. It fails closed and exits non-zero, so you can gate CI on it — and every subtask's cost is on a hash-chained ledger you can audit.
You can't buy more throughput. You can use all of yours.
On the API, the major vendors' published rate limits are now flat across self-serve tiers — Claude Fable 5 is 100K output tokens/minute per organization at every tier, roughly 30–50 concurrent agentic workers, and no plan upgrade raises it. The only compliant way past a single vendor's ceiling is fanning one task across accounts you already hold.
We never mark up a token — they run on your account at your provider's price. What we do is spend fewer of them, and reach ceilings you can't buy past:
Batch lanes (−50% on three of four majors), cache-aware routing (reads at 0.1× and exempt from your input limit), class routing (a 25–50× price spread across the pool), and multi-account fan-out across your own Bedrock, Vertex, Azure and direct orgs — same models, separate ceilings, one bill.
| lever | effect | on whose tokens |
|---|---|---|
| Batch lane | −50% | yours, at list |
| Cache-aware routing | 0.1× reads, ITPM-exempt | yours, at list |
| Function → class routing | 25–50× spread | yours, at list |
| Multi-account fan-out | N× the org ceiling | yours, at list |
Levers verified against Anthropic, OpenAI, Google and xAI official pricing and rate-limit docs (2026). Every run report shows list cost, actual cost, and which lever saved what — savings are a receipt, not a promise.
The part where we talk about your keys.
Encrypted, never shown
AES-256-GCM at rest under a server master key. The secret never appears in an API response, a log line, a run record, or a report.
Contained to your account
Keys serve only your runs. No share switch exists, and the scheduler's pool for a run is exactly your own active keys — provable, not asserted.
Ceilings you set
Rate windows and a hard USD cap per key. The scheduler treats them as physics; a key that hits its cap simply stops taking work.
The gate at the door
Keys whose own plan terms forbid backend use — coding-plan keys, subscription OAuth tokens — are refused at upload, with the provider's own clause quoted and a pointer to the right key.
An audit ledger
Every run's cost is attributed on a SHA-256 hash-chained ledger. Export it and verify the chain yourself — tampering is detectable by you, not just by us.
The honest bit
We can't make a token cheaper than list — nobody legally can. What we do is spend fewer of them, and every run proves the delta. No wallet, no cash-out, no token.
The name
A one-pager on what "Nubivola" means — and why it fits.
Nubivola
nūbēs (cloud) + volāre (to fly) — Latin, "cloud-flying"
Nūbivolus is a rare Latin poetic compound — the word for what moves through and above the clouds. Every AI vendor is a cloud — Anthropic, OpenAI, Google, xAI, DeepSeek — each with its own weather: rate limits, quotas, price fronts, outages. Most tools shelter inside a single cloud. Nubivola treats the clouds as terrain to fly over: your task is split into a graph, and each piece takes wing to whichever of your own accounts serves it best and cheapest at that moment. When one cloud storms — a 429, an outage, an exhausted budget — the work banks mid-flight to the next, without the run ever touching the ground.
FAQ
Can other people's jobs ever use my key?
No, structurally. Keys are contained to your account; the scheduler cannot see other accounts' keys, and there is no sharing feature to misconfigure. The run engine's pool is exactly your own active keys, and it's an invariant we test in CI.
Do you mark up tokens or resell capacity?
Never. Tokens run on your account at your provider's price. Our value is using fewer of them and reaching ceilings you can't buy past — not a margin on inference. That's also why there's nothing about us a CISO has to trust with your quota.
Is this renting out my Claude or ChatGPT subscription?
No. Subscriptions are per-human and off-limits — pooling them breaks every provider's terms, and we won't build it. Nubivola runs on API keys and cloud accounts, which are yours to use as you like. Coding-plan keys and subscription OAuth tokens are refused at the door.
What if a key hits a rate limit mid-run?
The key cools down pool-wide and the call retries on the next best key — which may belong to a different vendor, because the plan asks for a model class, not a model name. If every window you own is shut, the run pauses, burns no retries, and resumes when the first window rolls.
How do I know the output isn't garbage?
Deterministic gates run first, then an independent model that didn't write the work returns a structured verdict, then a bounded repair loop. Unverified output exits non-zero, so you can gate CI on it.
What does it cost?
Per seat, plus metered runs for CI-scale automation. Tokens are never marked up — our fee grows with your team, not your bill. Every run report shows what it cost and what naive routing would have cost.
Details & fine print
Your keys. AES-256-GCM encrypted at rest under a server master key; secrets never appear in API responses, logs, or run records. Every key is private to your account by construction — there is no sharing surface.
Compliance. Plan and subscription credentials whose terms ban backend use — coding-plan keys (sk-sp-…, tp-…), subscription OAuth tokens (sk-ant-oat…) — are refused at upload with a pointer to the provider's pay-as-you-go key. Consumer subscription seats (Claude Max, ChatGPT and the like) are per-human and out of scope; Nubivola runs on API keys and cloud accounts only.
Prototype status. Registration is open; bearer tokens are shown once and pasted back to enter the console. Real-provider mode needs working API keys with billing enabled; a built-in simulated provider exercises the full plan → shard → verify pipeline without spending a token. The pooled-intelligence campaign layer (many contributors, one goal, everyone on their own keys) is the next roadmap item — see the project docs.