Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • TRADING
  • SUBMIT
Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • TRADING
  • SUBMIT
Crypto Flexs
Home»ADOPTION NEWS»NVIDIA’s TensorRT-LLM MultiShot improves AllReduce performance with NVSwitch.
ADOPTION NEWS

NVIDIA’s TensorRT-LLM MultiShot improves AllReduce performance with NVSwitch.

By Crypto FlexsNovember 3, 20242 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
NVIDIA’s TensorRT-LLM MultiShot improves AllReduce performance with NVSwitch.
Share
Facebook Twitter LinkedIn Pinterest Email

just alvin
November 3, 2024 02:47

To improve multi-GPU communication efficiency, NVIDIA introduces TensorRT-LLM MultiShot, leveraging NVSwitch technology to perform AllReduce operations up to 3x faster.





NVIDIA has unveiled TensorRT-LLM MultiShot, a new protocol designed to improve the efficiency of multi-GPU communication, especially for generative AI workloads in production environments. According to NVIDIA, this innovation leverages NVLink switch technology to significantly increase communication speeds by up to 3x.

Challenges of existing AllReduce

Low-latency inference is critical in AI applications and often requires multi-GPU setups. However, the existing AllReduce algorithm, which is essential for GPU computation synchronization, may be inefficient as it involves multiple data exchange steps. Traditional ring-based approaches require 2N-2 steps. Here N is the number of GPUs, which increases latency and synchronization issues.

TensorRT-LLM multishot solution

TensorRT-LLM MultiShot solves these problems by reducing the latency of AllReduce operations. This leverages the multicast capabilities of NVSwitch to allow a GPU to send data to all other GPUs simultaneously with minimal communication steps. This results in only two synchronization steps being performed regardless of the number of GPUs involved, significantly increasing efficiency.

The process is divided into ReduceScatter tasks and AllGather tasks. Each GPU accumulates part of the resulting tensor and then broadcasts the accumulated result to all other GPUs. This method reduces per-GPU bandwidth and improves overall throughput.

Implications for AI Performance

Introducing TensorRT-LLM MultiShot can achieve nearly 3x speedup over existing methods, especially useful for scenarios that require low latency and high parallelism. These advancements allow for reduced latency or increased throughput at a given latency, potentially enabling ultra-linear scaling using more GPUs.

NVIDIA emphasizes the importance of understanding workload bottlenecks to optimize performance. The company is working closely with developers and researchers to implement new optimizations, with the goal of continuously improving the performance of the platform.

Image source: Shutterstock


Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Posts

AAVE Price Prediction: $100 is the wall. Factors that can destroy or bury a wall include:

July 25, 2026

Multicoin Capital has made its first Hyperliquid ecosystem investment in Trasia, an Asia-focused trading platform.

July 17, 2026

Polymarket Probability Price The probability that the United States will invade Iran before 2027 is 16.5%.

July 9, 2026
Add A Comment

Comments are closed.

Recent Posts

CT3 Begins Preparing Its Ecosystem for the Launch of the CT3GB Economy

August 10, 2026

Syntetika Launches Tokenization Hub Bringing Regulated Investment Strategies Onchain

August 10, 2026

Bitmine Immersion Technologies (BMNR) Announces ETH Holdings Reach 5.81 Million Tokens, and Total Crypto and Total Cash Holdings of $11.6 Billion

August 10, 2026

MEXC Sponsors Yohani’s Colombo Concert, Bridging Sri Lankan Culture and Global Digital Finance

August 10, 2026

Beyond the Headline Bonus -How to Measure Real Value at a Crypto Casino

August 8, 2026

Bybit Sues North Korea and Lazarus Group, Secures Preliminary Injunction Freezing Stolen Assets in Landmark Crypto Asset Recovery Effort

August 8, 2026

Carbon Launches TradFi-Native On-Chain Derivatives Venue With 950+ Markets in One Account

August 7, 2026

MEXC Lists New Ondo Tokenized Stock Pairs Spanning AI Infrastructure, Semiconductor and Rare Earth Sectors

August 7, 2026

ORBS) Reports Total Holdings of Approximately $378 Million, Includes OpenAI, Beast Industries, More Than 16,000 ETH and Nearly 302 Million WLD Tokens

August 6, 2026

ChangeNOW Brings Martin Masser Into Its Crypto Super App

August 5, 2026

MEXC 0808 debuts as an annual brand event with Stock Season and a $500,000 prize pool

August 5, 2026

Crypto Flexs is a Professional Cryptocurrency News Platform. Here we will provide you only interesting content, which you will like very much. We’re dedicated to providing you the best of Cryptocurrency. We hope you enjoy our Cryptocurrency News as much as we enjoy offering them to you.

Contact Us : Partner(@)Cryptoflexs.com

Top Insights

CT3 Begins Preparing Its Ecosystem for the Launch of the CT3GB Economy

August 10, 2026

Syntetika Launches Tokenization Hub Bringing Regulated Investment Strategies Onchain

August 10, 2026

Bitmine Immersion Technologies (BMNR) Announces ETH Holdings Reach 5.81 Million Tokens, and Total Crypto and Total Cash Holdings of $11.6 Billion

August 10, 2026
Most Popular

These two altcoins can be listed on Coinbase

February 27, 2024

BVI court freezes $1 billion assets of Three Arrows Capital founder

December 21, 2023

Mantle and VeChain are bullish. Investors Accumulating Pullix

February 1, 2024
  • Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms and Conditions
© 2026 Crypto Flexs

Type above and press Enter to search. Press Esc to cancel.