Synopsys and Ansys power the future of innovation—connecting silicon to systems.
Ansys potenzia la nuova generazione di ingegneri
Gli studenti hanno accesso gratuito a software di simulazione di livello mondiale.
Progetta il tuo futuro
Connettiti a Ansys per scoprire come la simulazione può potenziare la tua prossima innovazione.
Gli studenti hanno accesso gratuito a software di simulazione di livello mondiale.
Connettiti a Ansys per scoprire come la simulazione può potenziare la tua prossima innovazione.
For users of Ansys Fluent fluid simulation software who are evaluating new graphics processing unit (GPU) hardware, the key question is not simply which GPU is fastest. It is which platform delivers the best return on investment (ROI), balancing simulation turnaround time, hardware cost, and overall engineering productivity.
Across a range of benchmark cases, both the NVIDIA DGX B200 and RTX PRO 6000 Blackwell Workstation GPU demonstrate strong ROI through significant reductions in simulation time compared to CPU-based configurations. This acceleration not only improves efficiency, but also enables engineers to run larger, higher-fidelity models within similar timelines, increasing confidence in simulation results and supporting better design decisions.
The latest Fluent 2026 R1 results further clarify these trade-offs. With new benchmark data on NVIDIA DGX B200 and RTX PRO 6000 Blackwell systems, users can better understand how server-class GPUs fit into Fluent workflows. Public Ansys benchmark results already include Fluent 2026 R1 performance for the B200, while the RTX PRO 6000 Blackwell emerges as a strong and accessible option for smaller deployments.
GPU performance in computer-aided engineering (CAE) applications such as Fluent is inherently case dependent. The performance achieved on a given GPU is determined not only by the hardware itself, but also by key characteristics of the computational fluid dynamics (CFD) simulation, including model size (cell count), physics complexity, and solver type.
In this context, it is important to clearly distinguish between smaller and larger simulations. Smaller CFD simulations — typically below ~10 million cells — are less computationally intensive and require fewer compute resources. As a result, they tend to benefit less from scaling across multiple high-end GPUs. Larger simulations, by contrast, place significantly higher demands on compute, memory capacity, and memory bandwidth, making them much better suited to multi-GPU acceleration.
This is also reflected in typical CPU baselines used for Fluent workloads:
As model size increases, the required CPU infrastructure grows rapidly, which is precisely where GPU acceleration begins to provide the strongest value proposition.
The Fluent benchmarks on NVIDIA DGX B200 systems illustrate this behavior clearly. Larger cases such as the DrivAer_250M car model show strong scaling as GPU count increases, while smaller cases like combustor_24M, a model simulating internal combustion, exhibit more limited but still significant gains. This behavior is consistent with Amdahl's Law: As simulation size increases, a larger fraction of the workload can be executed in parallel, enabling more efficient scaling across multiple GPUs. Smaller models, by contrast, encounter diminishing returns sooner because serial operations and communication overhead represent a greater share of total runtime.
Equally important, these results must be interpreted on a cost-equivalent basis rather than by comparing raw hardware performance alone. NVIDIA DGX B200 and RTX PRO 6000 Blackwell are positioned at different points on the performance–cost curve. When normalized for comparable investment, the benchmarks highlight a key trade-off:
For example, in the below DrivAer_250M benchmark, multi-GPU scaling enables substantial performance gains on an 8-way NVIDIA DGX B200 supercomputer at comparable cost relative to a 6-node, 1032-CPU core baseline system.
Relative speedup for the DrivAer_250m case, demonstrating strong multi-GPU scaling and substantial throughput gains (up to ~8X) at a cost-comparable configuration versus CPU baseline system
For the DrivAer_50M case, it is particularly important to note that the CPU baseline configuration (~516 cores) and a 4x RTX PRO 6000 Blackwell server are roughly equivalent in overall system cost. When compared on this like-for-like basis, the GPU configuration delivers approximately a ~4.6X speedup over the CPU baseline, demonstrating a clear performance-per-dollar advantage even for mid-sized simulations.
By contrast, increasing investment toward a 4x B200 GPU configuration (at roughly ~30% higher system cost) further boosts performance to ~7.5X, illustrating how incremental spend on higher-end GPU systems translates directly into higher throughput.
The key takeaway is not simply identifying the fastest GPU, but understanding which configuration delivers the best performance per dollar for a given simulation profile. High-end GPU systems like the DGX B200 are optimized for large-scale, throughput-driven environments, while GPUs such as the RTX PRO 6000 Blackwell provide a more accessible and cost-efficient option for smaller or mixed workloads.
Relative performance for the DrivAer_50m case, showing ~4.6X speedup from an equivalent-cost GPU configuration versus CPU baseline and scaling up to ~7.5X with higher-end DGX B200 — highlighting strong performance-per-dollar gains
The examples above primarily represent external aerodynamics and segregated solver workflows. Two additional benchmark cases are shown below to extend coverage to additional physics domains (e.g., combustion), smaller model sizes, and the pressure-based coupled solver.
Relative speedup of the combustor_24m case, extending coverage to internal flow physics and showing strong multi-GPU scaling with consistent performance gains of NVIDIA DGX B200 over RTX Pro 6000 Blackwell and CPU baseline systems
Relative speedup of the exhaust_system_33m case, showing efficient multi-GPU scaling and consistent performance gains of NVIDIA DGX B200 platform over RTX PRO 6000 Blackwell and CPU baseline systems
The B200 results reinforce the value of high-end GPUs for large-production CFD workloads. This is where high memory bandwidth and multi-GPU scaling matter most. Each NVIDIA DGX B200 GPU has 192GB of memory and 8TB/s memory bandwidth, which helps explain why it is suited to larger Fluent models and throughput-oriented environments.
For organizations running very large models or managing multiple simulations in parallel, B200-class systems are particularly well-suited to maximizing throughput and overall cost efficiency across simulation portfolios.
The RTX PRO 6000 Blackwell is a different proposition. This GPU has 96GB of memory and 1,792GB/s memory bandwidth. These GPUs support visualization workflows while server GPUs typically offer higher peak compute performance and memory headroom.
That makes the RTX PRO 6000 relevant for Fluent users who want a high-end workstation focused on the full simulation workflow rather than dedicated server infrastructure. For many users, that is an attractive middle ground: substantial GPU acceleration in a single-box system that also supports the broader simulation workflow.
GPU selection should start with practical constraints rather than headline performance numbers.
Approximate GPU memory per 1 million fluid cells with a two-equation turbulence model and active energy equation
Beyond hardware choice, a few Fluent launch and run settings can improve overall efficiency. For maximum GPU solver speed, we recommend using one central processing unit (CPU) core per GPU card. When faster case read and write or post-processing matters, using “gpu_remap” with more CPU cores can improve input/output (I/O) and graphical user interface (GUI) responsiveness, even if the GPU solve itself is slightly slower.
It is also worth minimizing unnecessary communication between the GPU solver and the CPU or GUI. Less frequent reporting reduces data transfers, and “gpu_async” can enable calculations to continue while data is passed back for monitors and plots. For users balancing memory and accuracy, “gpu_hybrid_precision” can provide nearly double-precision accuracy with lower RAM demand and reduced computation time.
The broader value of Fluent software on NVIDIA GPUs is already visible in customer stories.
The latest Fluent 2026 R1 results on the DGX B200 GPU and RTX PRO 6000 Blackwell are useful because they help Fluent users match hardware class to workload.
For larger-production CFD jobs and multi-GPU scaling, DGX B200-class systems are the stronger fit. For users running smaller jobs, the RTX PRO 6000 Blackwell is likely to be more relevant. The right answer depends on the model size, usage pattern, and budget at least as much as on raw benchmark numbers.
Interested in testing the Fluent GPU Solver for yourself? Request a 30-day free trial.
The Ansys Advantage blog, featuring contributions from Ansys and other technology experts, keeps you updated on how Ansys simulation is powering innovation that drives human advancement.
Se devi affrontare sfide di progettazione, il nostro team è a tua disposizione per assisterti. Con una vasta esperienza e un impegno per l'innovazione, ti invitiamo a contattarci. Collaboriamo per trasformare i tuoi ostacoli ingegneristici in opportunità di crescita e successo. Contattaci oggi stesso per iniziare la conversazione.