Skip to Main Content
Countries & Regions

Synopsys and Ansys power the future of innovation—connecting silicon to systems.

Which NVIDIA Blackwell GPU Delivers the Best Fluent ROI?

July 22, 2026

READ ALOUD

PAUSE READ

Wim Slagter | Partnerships, Senior Director, Ansys, part of Synopsys
fluent-gpu-banner

For users of Ansys Fluent fluid simulation software who are evaluating new graphics processing unit (GPU) hardware, the key question is not simply which GPU is fastest. It is which platform delivers the best return on investment (ROI), balancing simulation turnaround time, hardware cost, and overall engineering productivity.

Across a range of benchmark cases, both the NVIDIA DGX B200 and RTX PRO 6000 Blackwell Workstation GPU demonstrate strong ROI through significant reductions in simulation time compared to CPU-based configurations. This acceleration not only improves efficiency, but also enables engineers to run larger, higher-fidelity models within similar timelines, increasing confidence in simulation results and supporting better design decisions.

The latest Fluent 2026 R1 results further clarify these trade-offs. With new benchmark data on NVIDIA DGX B200 and RTX PRO 6000 Blackwell systems, users can better understand how server-class GPUs fit into Fluent workflows. Public Ansys benchmark results already include Fluent 2026 R1 performance for the B200, while the RTX PRO 6000 Blackwell emerges as a strong and accessible option for smaller deployments.

Why These Results Matter

GPU performance in computer-aided engineering (CAE) applications such as Fluent is inherently case dependent. The performance achieved on a given GPU is determined not only by the hardware itself, but also by key characteristics of the computational fluid dynamics (CFD) simulation, including model size (cell count), physics complexity, and solver type.

In this context, it is important to clearly distinguish between smaller and larger simulations. Smaller CFD simulations — typically below ~10 million cells — are less computationally intensive and require fewer compute resources. As a result, they tend to benefit less from scaling across multiple high-end GPUs. Larger simulations, by contrast, place significantly higher demands on compute, memory capacity, and memory bandwidth, making them much better suited to multi-GPU acceleration.

This is also reflected in typical CPU baselines used for Fluent workloads:

  • <10M cells: ~172 CPU cores (1 node)
  • 10–45M cells: ~344 CPU cores (2 nodes)
  • 45–100M cells: ~516 CPU cores (3 nodes)
  • >100M cells: ~1032 CPU cores (6 nodes)

As model size increases, the required CPU infrastructure grows rapidly, which is precisely where GPU acceleration begins to provide the strongest value proposition.

The Fluent benchmarks on NVIDIA DGX B200 systems illustrate this behavior clearly. Larger cases such as the DrivAer_250M car model show strong scaling as GPU count increases, while smaller cases like combustor_24M, a model simulating internal combustion, exhibit more limited but still significant gains. This behavior is consistent with Amdahl's Law: As simulation size increases, a larger fraction of the workload can be executed in parallel, enabling more efficient scaling across multiple GPUs. Smaller models, by contrast, encounter diminishing returns sooner because serial operations and communication overhead represent a greater share of total runtime. 

Equally important, these results must be interpreted on a cost-equivalent basis rather than by comparing raw hardware performance alone. NVIDIA DGX B200  and RTX PRO 6000 Blackwell are positioned at different points on the performance–cost curve. When normalized for comparable investment, the benchmarks highlight a key trade-off:

  • High-end, multi-GPU systems deliver significantly higher throughput for large, compute-intensive simulations
  • More accessible configurations provide strong acceleration for smaller or moderately sized workloads, with lower cost

For example, in the below DrivAer_250M benchmark, multi-GPU scaling enables substantial performance gains on an 8-way NVIDIA DGX B200 supercomputer  at comparable cost relative to a 6-node, 1032-CPU core baseline system.

1-drivaer-250m-rev.jpg

Relative speedup for the DrivAer_250m case, demonstrating strong multi-GPU scaling and substantial throughput gains (up to ~8X) at a cost-comparable configuration versus CPU baseline system

For the DrivAer_50M case, it is particularly important to note that the CPU baseline configuration (~516 cores) and a 4x RTX PRO 6000 Blackwell server are roughly equivalent in overall system cost. When compared on this like-for-like basis, the GPU configuration delivers approximately a ~4.6X speedup over the CPU baseline, demonstrating a clear performance-per-dollar advantage even for mid-sized simulations.

By contrast, increasing investment toward a 4x B200 GPU configuration (at roughly ~30% higher system cost) further boosts performance to ~7.5X, illustrating how incremental spend on higher-end GPU systems translates directly into higher throughput.

The key takeaway is not simply identifying the fastest GPU, but understanding which configuration delivers the best performance per dollar for a given simulation profile. High-end GPU systems like the DGX B200 are optimized for large-scale, throughput-driven environments, while GPUs such as the RTX PRO 6000 Blackwell provide a more accessible and cost-efficient option for smaller or mixed workloads.

2-drivaer-50m-rev.jpg

Relative performance for the DrivAer_50m case, showing ~4.6X speedup from an equivalent-cost GPU configuration versus CPU baseline and scaling up to ~7.5X with higher-end DGX B200 — highlighting strong performance-per-dollar gains

The examples above primarily represent external aerodynamics and segregated solver workflows. Two additional benchmark cases are shown below to extend coverage to additional physics domains (e.g., combustion), smaller model sizes, and the pressure-based coupled solver.

3-cumbustor-24m

Relative speedup of the combustor_24m case, extending coverage to internal flow physics and showing strong multi-GPU scaling with consistent performance gains of NVIDIA DGX B200 over RTX Pro 6000 Blackwell and CPU baseline systems

4-exhaust-33m

Relative speedup of the exhaust_system_33m case, showing efficient multi-GPU scaling and consistent performance gains of NVIDIA DGX B200 platform over RTX PRO 6000 Blackwell and CPU baseline systems

What the NVIDIA DGX B200 Results Suggest

The B200 results reinforce the value of high-end GPUs for large-production CFD workloads. This is where high memory bandwidth and multi-GPU scaling matter most. Each NVIDIA DGX B200 GPU has 192GB of memory and 8TB/s memory bandwidth, which helps explain why it is suited to larger Fluent models and throughput-oriented environments.

For organizations running very large models or managing multiple simulations in parallel, B200-class systems are particularly well-suited to maximizing throughput and overall cost efficiency across simulation portfolios.

Where the RTX PRO 6000 Fits

The RTX PRO 6000 Blackwell is a different proposition. This GPU has 96GB of memory and 1,792GB/s memory bandwidth. These GPUs support visualization workflows while server GPUs typically offer higher peak compute performance and memory headroom.

That makes the RTX PRO 6000 relevant for Fluent users who want a high-end workstation focused on the full simulation workflow rather than dedicated server infrastructure. For many users, that is an attractive middle ground: substantial GPU acceleration in a single-box system that also supports the broader simulation workflow.

Hardware Considerations That Matter

GPU selection should start with practical constraints rather than headline performance numbers.

  • First, memory capacity is critical. If the model does not fit, multiple GPUs should be used, but total available memory still has to cover the model and overhead. We also recommend having at least as much system random-access memory (RAM) as total GPU memory — and potentially more for some mesh types.
5-mesh-type-table

Approximate GPU memory per 1 million fluid cells with a two-equation turbulence model and active energy equation

  • Second, memory bandwidth matters a great deal for CFD software. In most cases, Fluent users also benefit from higher memory bandwidth, which is one reason that new server GPUs are suitable for different types of workloads.
  • Third, the deployment model matters. A workstation GPU may be the better fit for flexibility and accessibility while a server GPU is the stronger option for very large jobs and sustained multi-GPU use.
  • Fourth, software compatibility should not be overlooked. Fluent 2025 R2 and 2026 R1 are supported with CUDA 12.8. Beginning with Fluent 2026 R1 SP02, PyUDFs support CUDA 12.8 as well as newer CUDA releases.

Software Settings That Can Improve Efficiency

Beyond hardware choice, a few Fluent launch and run settings can improve overall efficiency. For maximum GPU solver speed, we recommend using one central processing unit (CPU) core per GPU card. When faster case read and write or post-processing matters, using “gpu_remap” with more CPU cores can improve input/output (I/O) and graphical user interface (GUI) responsiveness, even if the GPU solve itself is slightly slower.

It is also worth minimizing unnecessary communication between the GPU solver and the CPU or GUI. Less frequent reporting reduces data transfers, and “gpu_async” can enable calculations to continue while data is passed back for monitors and plots. For users balancing memory and accuracy, “gpu_hybrid_precision” can provide nearly double-precision accuracy with lower RAM demand and reduced computation time.

Customer Examples

The broader value of Fluent software on NVIDIA GPUs is already visible in customer stories.

  • At Leonardo Helicopters, Fluent software on GPUs enabled models to run 2.6 times faster while using only one-third of the hardware resources, with estimated energy use reduced from 85 kWh to 15 kWh.
  • At Seagate, Fluent software on GPUs helped dramatically reduce runtimes and shorten design cycles. The blog highlights a 50-times-acceleration example.
  • At Honda, Fluent software on GPUs helped speed simulation solve times by 34X compared to running the same simulation on 1,920 cloud-based CPU cores, with a 38X cost reduction in compute expenses.

Finding the Right GPU

The latest Fluent 2026 R1 results on the DGX B200 GPU and RTX PRO 6000 Blackwell are useful because they help Fluent users match hardware class to workload.

For larger-production CFD jobs and multi-GPU scaling, DGX B200-class systems are the stronger fit. For users running smaller jobs, the RTX PRO 6000 Blackwell is likely to be more relevant. The right answer depends on the model size, usage pattern, and budget at least as much as on raw benchmark numbers.

Interested in testing the Fluent GPU Solver for yourself? Request a 30-day free trial.


Just for you. We have some additional resources you may enjoy.

TAKE A LOOK


wim-slagter.jpg
Senior Director, Partner Programs

Wim Slagter

Wim is senior director, partner programs at Ansys, part of Synopsys. In his role, he is responsible for the overall design and execution of the partner programs (including high-performance computing programs) within Corporate Development and Global Partnerships at Ansys. Wim has 30 years of experience in the business of engineering simulation software with management positions in software development, consulting, sales, and product management. Wim holds a doctorate degree in Aerospace Engineering from the Technical University of Delft in the Netherlands.

Recommendations

Which NVIDIA Blackwell GPU Delivers the Best Fluent ROI?

Which NVIDIA Blackwell GPU Delivers the Best Fluent ROI?

The latest Fluent 2026 R1 results on the DGX B200 GPU and RTX PRO 6000 Blackwell are useful because they help Fluent users match hardware class to workload.

The Science Behind Cycling

The Science Behind Cycling

As the 113th Tour de France gets ready to set off, researchers, innovators, and professional cycling team members give their technology perspectives.

King’s College London Reduces Congenital Heart Defect Diagnostic Risks With Neonatal Digital Twins

King’s College London Reduces Congenital Heart Defect Diagnostic Risks With Neonatal Digital Twins

Learn how King's College London uses simulation to revolutionize how clinicians predict and manage congenital heart disease (CHD) in unborn babies.

The Advantage Blog

The Ansys Advantage blog, featuring contributions from Ansys and other technology experts, keeps you updated on how Ansys simulation is powering innovation that drives human advancement.