CarbonForge containers
LLM inference optimisation
Company describes serving images built for a specific combination of hardware, model and workload, with a scheduler and a controller that adapts GPU operating parameters, running alongside vLLM in the customer's cloud and licensed by GPU class.
Learn more