TibiOS PRO AI Network

Turns a building into Frontier AI infrastructure. (advanced alternative)

TibiOS Private AI Network is an alternative mode for users who need to take AI usage far beyond a conventional experience.

It is designed for companies, professional teams and demanding user communities that need to work intensively with Frontier models and want the highest possible inference speed and capacity, without depending solely on a single machine or on cloud services.

TibiOS connects multiple TibiBoxes through a private, high-speed network and distributes the Experts of MoE models across the different nodes. This way, the infrastructure can grow by adding compute and memory capacity wherever it's needed.

It isn't meant to replace the usual TibiOS experience: it's an advanced mode for those who need higher speed and intensive use.

More capacity. More speed. More AI.
A new way to build private AI infrastructure for the most demanding users.

TibiOS Private AI Network: connects houses and a building into one distributed AI cluster, sharding MoE model experts across Orin nodes over a high-speed switch

TibiOS lets you connect houses, offices and entire buildings into a private computing network, turning many small GPUs into a single distributed AI Cluster. Instead of needing one giant GPU to run a Frontier model, TibiOS distributes its MoE model Experts across the different nodes and uses a high-speed switch to coordinate inference traffic. More nodes → more AI capacity. The infrastructure can grow from a few machines in a building to hundreds of distributed GPUs, while models and their data stay inside a private network.

The idea: your building becomes a distributed GPU. TibiOS automatically manages routing, placement, memory, load and communication between nodes, keeping the most-used Experts close to where they're needed. The result is a private infrastructure capable of offering access to Frontier models, high inference speed and virtually unlimited capacity by adding new nodes, without depending on a single, enormously expensive GPU.

One network. Many GPUs. One AI cluster.

TibiOS PRO AI Network — Technical Detail

TibiOS Private AI Network technical architecture: TibiOS AI Cluster Engine routing tokens across a node grid with Trunk, Experts and Others memory, plus the TibiOS Control Plane

Expert Sharding

In an MoE model, not every parameter takes part in each token. A router selects a small subset of Experts for every token. TibiOS takes advantage of this property by distributing the Experts across the different nodes in the cluster. Each node holds a shard of the Expert set, and the system dynamically determines where each selected Expert lives.

Token Router Expert IDs Node IDs Network Route Expert Execution Result

This way, the model's total capacity can be far greater than the memory of any single node.

TibiOS AI Cluster Engine

The TibiOS AI Cluster Engine acts as the cluster's intelligent coordination layer. It keeps a global view of the available resources and makes routing, scheduling and placement decisions. Among other parameters, the scheduler can be aware of:

  • GPU and compute capacity of each node
  • Available memory
  • Experts resident on each node
  • Expert cache state
  • Current GPU load
  • Network topology
  • Available bandwidth
  • Latency between nodes
  • Historical routing locality

This lets Expert placement account for compute, memory and communication cost at the same time.

Trunk Residency

The model's trunk has a different access pattern than the Experts: every token walks through its layers in the same order, so a conventional LRU cache isn't well suited to managing a trunk that doesn't fully fit in memory. TibiOS uses a fixed-pinning model whenever possible — resident layers stay in memory, while the rest can use secondary storage as a streaming mechanism, so the trunk doesn't become a constant per-token I/O stream. When the model and the hardware allow it, the ideal architecture keeps the trunk fully resident in RAM/VRAM and reserves storage for non-resident Experts.

Expert Cache

Unlike the trunk, Experts can benefit from caching mechanisms because their usage depends on routing. Frequently used Experts can stay resident in memory, while others live on secondary storage and load on demand. The effectiveness of this mechanism depends fundamentally on routing locality: the higher the probability that upcoming tokens reuse already-resident Experts, the higher the hit rate and the lower the storage traffic. That's why TibiOS treats Expert locality as a core variable for distributed scheduling.

Network-Aware Scheduling

In a distributed architecture, the network is part of the inference system. TibiOS can jointly consider:

  • Compute cost
  • Memory capacity
  • Expert locality
  • Bandwidth
  • RTT (round-trip time)
  • Node load
  • Proximity within the network topology

This avoids placements that generate unnecessary traffic and favors co-locating Experts with related usage patterns within the same network domain. The infrastructure can use high-speed Ethernet — 10, 25 or 100 GbE — depending on scale and hardware.

Communication and Batching

The main challenge of a distributed MoE architecture is preventing inter-node communication from dominating inference time. TibiOS is designed to group operations headed to the same node and process them through batching whenever possible. Instead of performing multiple individual communications per token, the system can group tokens headed to the same Experts and minimize the number of network exchanges. This amortizes latency and increases effective GPU utilization.

Horizontal Scalability

Cluster capacity grows by adding nodes. Each new node simultaneously contributes:

  • GPU capacity
  • Memory
  • Expert cache capacity
  • Storage
  • Aggregate inference capacity

However, TibiOS does not assume linear scalability. The scheduler's goal is to maximize scaling efficiency, accounting for both compute and communication. The key metric is the ratio between the throughput actually achieved and the theoretical throughput expected when adding new nodes.

Private AI Fabric

The architecture can be deployed inside a home, an office, a building, a campus or a private infrastructure. Nodes form a distributed pool of AI resources, while TibiOS abstracts away the complexity of deciding where Experts live, where to run each operation, and how to use the available resources. The infrastructure can grow progressively from a handful of nodes to clusters of dozens or hundreds of GPUs.

The TibiOS Architecture

TibiOS's core proposition isn't simply connecting GPUs. It's building an infrastructure layer capable of treating compute, memory, Experts and network as a single distributed resource. This enables a new way to scale inference: instead of buying an ever-larger GPU, add small nodes to a private network and intelligently distribute the model's capacity across them. The end goal is to turn many small GPUs into a single logical inference platform, capable of running Frontier models and increasing aggregate throughput by adding new nodes.

TibiOS Private Network

A high-speed private network for distributed AI

TibiOS Private Network makes it possible to deploy a private network infrastructure that connects TibiBoxes and other TibiOS nodes with low latency and high bandwidth, without interfering with each user's regular Internet connection.

The architecture separates two types of traffic:

  • Internet traffic: keeps using the user's router and connection as usual.
  • AI traffic: flows over a private network optimized for TibiOS, dedicated to node-to-node communication, distributed inference, Expert caches and future distributed model deployments.

This separation makes it possible to build an AI network inside a building, home, office, campus or multi-apartment installation.

TibiOS Private Network installation: Floor Switch distributing 2.5GbE to apartments over SFP+ uplinks to the Core, and a TibiOS Mini Switch inside each apartment separating router/Internet traffic from the TibiBox AI node

How it works

The infrastructure is split into two main tiers:

1. TibiOS Floor Switch

The Floor Switch distributes the high-speed network from the Core to the different apartments on a floor. Each apartment can have a dedicated 2.5GbE connection to the floor switch.

The Floor Switch includes:

  • 8 × 2.5GbE RJ45 ports for the apartments
  • 2 × SFP+ 10/25GbE uplinks to the Core
  • VLANs for network segmentation
  • QoS to prioritize AI traffic
  • Remote management and monitoring
  • High-speed connectivity between the infrastructure's different zones

Cabling can be run with Cat6A, Cat7 or fiber, depending on the design and distances involved.

As a reference, eight 2.5GbE connections represent up to 20Gbps of aggregate traffic toward the Floor Switch, though actual throughput depends on the uplinks, network configuration and traffic pattern.

2. TibiOS Mini Switch inside the apartment

The connection coming from the Floor Switch reaches the apartment over high-speed Ethernet. Inside the apartment, a TibiOS Mini Switch is installed, acting as the separation point between the user's conventional network and the private AI network.

The Mini Switch can provide:

  • 3–5 × 2.5GbE ports
  • VLAN and/or QoS
  • Separation between Internet traffic and AI traffic
  • Plug & play connectivity
  • One port toward the apartment's router
  • One or more ports toward TibiBoxes and TibiOS nodes

This way, the user can keep using their router, Wi-Fi, Internet services and personal devices as normal, while TibiOS uses an independent private network for node-to-node communication.

Installation guide

  1. Step 1 — Connect the Floor Switch to the Core. Install the TibiOS Floor Switch in the floor's cabinet or distribution point. Connect its SFP+ uplinks to the Core Switch using 10GbE, 25GbE or whatever speed the infrastructure supports. Larger installations may use redundant links or specific aggregation setups depending on the network design.
  2. Step 2 — Connect the apartments. Connect the Floor Switch's 2.5GbE ports to the apartments through the building's structured cabling (Floor Switch → Apartment 1, 2, 3… up to 8). The infrastructure can use Cat6A/Cat7 or fiber depending on the building's needs.
  3. Step 3 — Install the TibiOS Mini Switch. Inside each apartment, connect the cable coming from the Floor Switch to the Mini Switch's input port. The Mini Switch becomes the local distribution point for the TibiOS network.
  4. Step 4 — Connect the apartment's router. Connect one Mini Switch output to the user's router, which keeps providing Internet, Wi-Fi, home services and personal-device connectivity as usual — the user doesn't need to change how they use the Internet.
  5. Step 5 — Connect the TibiBox. Connect another Mini Switch port to the TibiBox, which acts as a node on the private TibiOS network and can include an NVIDIA GPU, a TibiOS Node, the inference engine, the Expert Cache and other distributed AI services. TibiOS nodes connected to the private infrastructure can communicate with each other without using the public Internet connection.

Network segmentation

The logical separation between networks is done through VLANs and QoS policies, depending on the deployment configuration. The architecture keeps two domains apart: the Internet domain (the user's router, Wi-Fi and conventional services) and the TibiOS Private Network domain (communication between TibiBoxes, inference nodes and other AI infrastructure components).

This prevents a download, a stream or any other regular Internet activity from unnecessarily consuming the bandwidth reserved for communication between TibiOS nodes.

Advantages

  • High speed: the infrastructure can use 2.5GbE, 10GbE, 25GbE links or higher, enabling networks suited for distributed AI workloads.
  • Low latency: traffic between nodes can stay within the private network, without depending on the Internet for internal communication.
  • Traffic separation: Internet and AI traffic can operate independently.
  • Scalability: the architecture can grow from a single apartment to multiple floors and multiple TibiBoxes.
  • Built for distributed AI: the network provides the infrastructure needed to connect TibiOS nodes participating in distributed workloads, including inference, node-to-node communication and model or Expert distribution.