Flexible inference + fine-tune framework
Put every tokenon the right hardware.
A flexible framework for experiencing heterogeneous LLM inference and fine-tune optimizations.
Illustrative heterogeneous routePipeline active
Compute path AAccelerator
Compute path BHost memory
Workload placementIllustrative
acceleratorhost
One model / multiple compute pathsFlexible by design