THE NOETON STACK · L4 / L5

L4 Network Fabric · Tiered Cognition

NETWORK FABRIC · TIERED COGNITION

Full capability comes from networked weight pools.

The Noeton does not pretend a single terminal can hold all of intelligence. L4 extends the machine into the network: a small on-device model gives millisecond responses and an offline floor, edge nodes host regional weight pools, and the cloud serves the largest models. A scheduler routes each request to the cheapest tier that satisfies it — a single request may even place prefill and decode on different tiers.

This is the Noeton's most controversial layer: network-mandatory design means an offline Noeton is deliberately limited. We believe capability density is worth that price — but the paper devotes a full section to the costs in autonomy, privacy, and equity rather than waving them away. Pull the offline switch in the tiered-cognition lab and feel the trade-off yourself.

On-device small model

Dictation, local search, basic control — the degraded-mode floor

Edge weight pool

Regionally shared mid-size models, tens of milliseconds away

Cloud full pool

The largest models and the full toolset — the capability ceiling

Split router

Dispatches request phases across tiers by compute/bandwidth profile

INTERACTIVE · LIVE

"Cheapest-first" is a bill, not a slogan. Switch strategies and watch the same request stream fork in cost.

01

Cheapest tier wins

Scheduling asks one question: the lowest-cost tier that meets the quality bar.

02

Capability is the network

Upgrading the model is like switching waterworks — invisible at the tap.

03

Costs faced honestly

Capability density is bought with mandatory connectivity; the risks are discussed in the paper, not buried.

Related lab: device·edge·cloud router (with offline switch) →