Optimize Dynamo Deployment in Minutes with AI Configurator
Optimizing large language model (LLM) serving is complex. Which framework offers the best perf? How do you choose between aggregated vs. disaggregated serving? If disaggregated, what’s the best prefill/decode split? What kind of parallelism should you use for each worker? Exploring this space can take weeks.
AI Configurator simplifies this to just minutes. Use the CLI to input your model, hardware, traffic characteristics, and goals, and the tool intelligently searches for an optimal configuration. AIC combines kernel-level benchmarks on real silicon with powerful simulation tools to (1) accurately model thousands of LLM inference scenarios, (2) suggest the configuration that best meets your needs and (3) the profile deployment manifests to make Dynamo deployment easy.
