INFERENCE · LAB 03 / 04

Requests land on the cheapest tier.

TIERED COGNITION · DEVICE → EDGE → CLOUD

The device holds the floor, the edge eats the latency, the cloud serves full intelligence.

Not every request deserves to wake a hundred billion parameters. Setting an alarm is well within an on-device model; summarizing a web page suits a mid-size model at the edge; only drafting a cross-border contract needs the full weight pool in the cloud. Tiered cognition schedules by one sentence: a request is always resolved at the cheapest tier that satisfies it.

This also means the Noeton is deliberately network-mandatory: the device guarantees only a degraded-mode floor, and full capability comes from networked weight pools. An offline Noeton is intentionally limited — a deliberate trade of autonomy for capability density, and the paper devotes a full section to its costs.

LIVE LAB
📱 Device small model · the floor 📡 Edge mid model · regional pool ☁️ Cloud largest models · full pool
tier hit
round trip
energy cost

Fire a request and see which tier it lands on. Then pull the network switch and try again — edge and cloud go dark, complex requests bounce back, and only the device floor stays alive: that is degraded mode. (Demo numbers.)

01

Cheapest first

The target is not fastest or strongest but just-enough: the lowest-latency, lowest-energy tier that meets the quality bar.

02

Capability is the network

Weight pools live in the network, not the box — upgrading the model is like switching waterworks without changing the tap.

03

Degraded, with a floor

Offline, the device still does dictation, local search, basic control — deliberately limited, never zero. The costs are discussed honestly in the paper.