Flexible inference + fine-tune framework

Put every tokenon the right hardware.

A flexible framework for experiencing heterogeneous LLM inference and fine-tune optimizations.

Illustrative heterogeneous routePipeline active
Compute path AAccelerator
Compute path BHost memory
Workload placementIllustrative
L1
L2
L3
L4
L5
L6
L7
L8
acceleratorhost
One model / multiple compute pathsFlexible by design