Authors
Updated
7 Oct 2026Form Number
LP2487Introduction
This sizing tool helps you determine the hardware and infrastructure required to deploy large language models for a wide range of AI workloads. Whether you're building a chatbot, Retrieval-Augmented Generation (RAG) application, AI agent, code assistant, or another generative AI solution, this tool helps you estimate the hardware resources needed to meet your performance and scalability requirements.
The tool includes two sizing options:
- Calculator: Sizes GPU-based deployments for your model and workload.
- CPU Sizing: Estimates capacity and cost for running inference on CPU-only Lenovo servers.
Version 2.0. Click the Full Change History link to see what's new.
Getting started
Scroll down to use the tool. To get started with the GPU calculator:
- Select the Use Case
- Select the desired GPU Type
- Click the Calculate GPU Requirements button
The tool automatically populates recommended preset values for the model and workload parameters, including concurrency, input and output token lengths, precision, and other deployment settings, which you can adjust as needed. Based on these inputs, it analyzes your workload and recommends an optimal infrastructure configuration. Learn more.
Use the Subscribe to Updates link to get notified when the tool gets updated, and use the Feedback link to send us comments and suggestions.
Using the sizing tool
The tool includes two sizing options:
- Calculator: Sizes GPU-based deployments for your model and workload.
- CPU Sizing: Estimates capacity and cost for running inference on CPU-only Lenovo servers.
Using the Calculator tool
Get started with these steps:
- Select the Use Case
- Select the desired GPU model
- Click the Calculate GPU Requirements button
The tool automatically populates recommended preset values for the model and workload parameters, including concurrency, input and output token lengths, precision, and other deployment settings, which you can adjust as needed. Based on these inputs, it analyzes your workload and recommends an optimal infrastructure configuration.
Custom use cases and models:
- Custom use case: Select Custom as the use case and set your own values under Advanced Parameters.
- Custom model: Select Custom as the model and enter its HuggingFace model ID in the Custom Model Name field (for example, mistralai/Mixtral-8x7B-v0.1), then press Enter. The model details are filled in automatically and can be reviewed under Advanced Parameters.
The Calculator results include:
- Recommended GPU count and a detailed VRAM breakdown
- Server platform and a recommended Lenovo Hybrid AI platform
- CPU and system memory requirements
- Storage, network adapters, and switching infrastructure
- Estimated power usage
- Inference performance metrics, including Time to First Token (TTFT), Inter-Token Latency (ITL), end-to-end request latency, throughput, and maximum supported concurrent users
Results can be downloaded as a PDF report, enabling you to evaluate different deployment scenarios before provisioning hardware.
Using the CPU Sizing tool
For CPU-based inference, open the CPU Sizing tab:
- Select a business use case
- Enter the number of concurrent users, active hours per day, server lifecycle, and electricity cost
- Review the recommended Lenovo servers and CPUs
For each option, the tool shows the supported capacity and the cost per million tokens compared with running the same workload in the cloud.
Use this tool to compare infrastructure options, optimize resource utilization, and make informed decisions when designing scalable and cost-effective AI deployments on Lenovo infrastructure. All estimates are based on analytical models and benchmark data and should be used as guidance for capacity planning rather than guaranteed production performance.
Use the Subscribe to Updates link to get notified when the tool gets updated, and use the Feedback link to send us comments and suggestions.
Configure and Buy
Full Change History
Changes in Version 2.0, published October 6, 2026:
- CPU Sizing: New tab for sizing CPU and memory requirements
- Memory breakdown: Clearer pie chart showing how GPU memory is used, also included in the PDF report
- Hybrid platform recommendations: The tool now suggests a matching Lenovo Hybrid AI platform for your workload
- More GPUs and models: Added NVIDIA Blackwell B300 and RTX PRO Blackwell GPUs, plus newer popular models
- Power: Estimated power usage now shown in the results
- Updated interface: Redesigned PDF report, improved dark mode and general layout tidy up
- Better accuracy: Improved latency, throughput and memory estimates, including support for MoE models and longer context lengths
Course Detail
Employees Only Content
The content in this document with a is only visible to employees who are logged in. Logon using your Lenovo ITcode and password via Lenovo single-signon (SSO).
The author of the document has determined that this content is classified as Lenovo Internal and should not be normally be made available to people who are not employees or contractors. This includes partners, customers, and competitors. The reasons may vary and you should reach out to the authors of the document for clarification, if needed. Be cautious about sharing this content with others as it may contain sensitive information.
Any visitor to the Lenovo Press web site who is not logged on will not be able to see this employee-only content. This content is excluded from search engine indexes and will not appear in any search results.
For all users, including logged-in employees, this employee-only content does not appear in the PDF version of this document.
This functionality is cookie based. The web site will normally remember your login state between browser sessions, however, if you clear cookies at the end of a session or work in an Incognito/Private browser window, then you will need to log in each time.
If you have any questions about this feature of the Lenovo Press web, please email David Watts at dwatts@lenovo.com.