Distributed local AI

Run one larger local model across the computers you already own

When a workload does not fit comfortably on one machine, Rex Connector prepares an ordered GGUF pipeline across connected computers and gives you one place to launch it, inspect it, and work with it

  • One ordered model pipeline
  • Existing local computers
  • Capacity first, with honest speed reporting
Example prepared topology One coordinator · Three workers
Your model Prepared as one ordered worker chain

Illustrative topology, not a performance benchmark

The job to be done

The model you want can be larger than one machine's safe memory budget

A common answer is to buy a larger accelerator, rent a cloud endpoint, or abandon the model choice altogether. Rex Connector is built around a fourth option: prepare the workload across computers you already control

Without a cluster

One machine becomes the hard limit

The model, runtime state, and working context all compete for finite memory on the same computer

With Rex Connector

Prepare one visible stage chain

Use live hardware telemetry and the model tensor manifest to place consecutive stages across selected computers

ReusePut existing machines to work
ControlKeep the pipeline on your network
VisibilitySee placement, links, state, and logs

From spare capacity to one runtime

Three steps from connected computers to a local Agent Session

1

Connect the computers

Select the coordinator and workers that can participate on your physical or virtual local network

Rex records the real links instead of assuming that every worker can reach every other worker
2

Prepare the model

Choose a GGUF file and let Prepare calculate ordered stage ranges from exact weight bytes and live memory telemetry

The prepared order and direct or relayed boundaries become an explicit launch plan
3

Launch and work

Start the prepared runtime and use the same Agent Session while commands, tool results, state, and logs remain visible

Compatibility is validated per model, build, placement, and topology

Watch the idea in motion

See why local AI does not have to stop at one computer

The video streams directly from YouTube. It starts muted when most of the player enters your screen, pauses when you scroll away, and keeps the normal player controls available

Muted autoplay is required by modern browsers. Use the player controls to enable sound or replay from the beginning

Numbers, not adjectives

Every speed claim should come with the exact setup that produced it

Public cluster benchmark matrixEvidence collection in progress
01Model and quantizationExact GGUF identity and size
02Every computerCPU, GPU, VRAM, RAM, and OS
03Network pathDirect or relayed links and transport
04Time to first tokenCold and warm runs kept separate
05Prompt throughputMeasured prompt tokens per second
06Generation speedMeasured generated tokens per second
07Single-node baselineIncluded only when the same workload fits
08Build identityPackaged app and native worker hash
Publication gate

No customer throughput number is published before a repeatable packaged run

Internal diagnostics can find bugs, but they are not marketing evidence. This page will replace this notice with reproducible tables only after the exact model, placement, worker build, network, and Agent Session E2E all pass together

Performance on your hardware

See how your own model, computers, and network perform together

Rex Connector prepares the available hardware as one visible pipeline and keeps the measurements attached to the exact model, topology, placement, and worker build that produced them

Use the computers you already have

Combine available accelerator and system memory for supported workloads while keeping model execution under your control

Prepare the route for your setup

Placement accounts for the model, available memory, accelerators, node balance, and measured links between selected computers

Inspect results with full context

Time to first token, prompt throughput, generation speed, topology, and build identity stay visible instead of being reduced to one unexplained number

A different job, not another model picker

Choose the runtime around your actual constraint

Rex Connector adds a prepared multi-computer pipeline when one host is the capacity limit, while keeping the complete local workflow in one application

ApproachBest fitCapacity boundaryWhat it gives you
Single-computer local runnerThe workload fits one hostThat host's safe compute and memory budgetSimple local path, limited by one machine
Cloud inference APIYou want managed remote capacityProvider plan and service limitsRecurring service dependence and remote processing
Rex Connector clusterOne local workload needs several connected computersValidated model, topology, and combined placementOne prepared local pipeline across your computers

Where the idea earns its keep

Concrete situations, not a list of components

01

Reuse idle office computers

You have several capable machines, but no single one has the safe memory budget for the local workload you want to evaluate

02

Keep sensitive work on your network

You want model execution and project prompts on computers you control while still using a visible agent and tool workflow

03

Evaluate capacity before buying new hardware

You want measured evidence from an existing topology before deciding whether a larger single accelerator is worth the cost

Compatibility without hand-waving

A model family name is not enough

Rex Connector validates the exact GGUF, tensor layout, packaged worker build, stage placement, state handoff, sampled-token output, and real Agent Session path before calling a distributed configuration supported

1

Exact GGUF architecture and tensor binding

2

Matching worker protocol build and source patch hash

3

Native batched-prefill and retained-session parity

4

Real packaged multi-stage Agent Session E2E

The public compatibility matrix is intentionally withheld until configurations clear the complete gate

Start with the topology you need

One-time licenses for one computer or a local cluster

Choose the license size that matches the number of computers you want to connect

Secure checkout

Choose a license

Select the maximum cluster size, enter an email, and continue to checkout

Direct answers

Questions to answer before you build a cluster

Will three computers be faster than one GPU

Not automatically, and often not when an equally capable single GPU can hold the same workload. Rex Connector's cluster path is primarily about capacity and hardware reuse

Can I just add every machine's VRAM together

No. Weight placement, runtime state, host memory, accelerator memory, network boundaries, and platform-specific memory accounting are separate constraints handled during Prepare

Which model should I buy this for

Use the public compatibility and benchmark matrix for verified configurations, or validate your own GGUF and topology in Rex Connector

Does it need the internet

Model execution and prompts can remain on your computers. Downloads, checkout, activation, and periodic license revalidation require internet access

Does Wi-Fi work

Reachability alone does not guarantee useful performance. Link latency, bandwidth, stability, and direct versus relayed routing must be measured for the prepared topology

Is this just another local model interface

The interface is supporting infrastructure. The product bet is the prepared ordered pipeline that lets one local workload use several connected computers

Try the current desktop build

Bring the computers, then prepare the exact topology

Release availability is being loaded for Windows, Linux, and Apple Silicon macOS

Help shape the evidence

Bring your model, hardware, and hardest objection

Follow development, compare exact setups, report failures, and help turn the benchmark matrix into something buyers can trust