Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • TRADING
  • HACKING
  • SLOT
  • CASINO
  • SUBMIT
Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • TRADING
  • HACKING
  • SLOT
  • CASINO
  • SUBMIT
Crypto Flexs
Home»ADOPTION NEWS»NVIDIA’s TensorRT-LLM MultiShot improves AllReduce performance with NVSwitch.
ADOPTION NEWS

NVIDIA’s TensorRT-LLM MultiShot improves AllReduce performance with NVSwitch.

By Crypto FlexsNovember 3, 20242 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
NVIDIA’s TensorRT-LLM MultiShot improves AllReduce performance with NVSwitch.
Share
Facebook Twitter LinkedIn Pinterest Email

just alvin
November 3, 2024 02:47

To improve multi-GPU communication efficiency, NVIDIA introduces TensorRT-LLM MultiShot, leveraging NVSwitch technology to perform AllReduce operations up to 3x faster.





NVIDIA has unveiled TensorRT-LLM MultiShot, a new protocol designed to improve the efficiency of multi-GPU communication, especially for generative AI workloads in production environments. According to NVIDIA, this innovation leverages NVLink switch technology to significantly increase communication speeds by up to 3x.

Challenges of existing AllReduce

Low-latency inference is critical in AI applications and often requires multi-GPU setups. However, the existing AllReduce algorithm, which is essential for GPU computation synchronization, may be inefficient as it involves multiple data exchange steps. Traditional ring-based approaches require 2N-2 steps. Here N is the number of GPUs, which increases latency and synchronization issues.

TensorRT-LLM multishot solution

TensorRT-LLM MultiShot solves these problems by reducing the latency of AllReduce operations. This leverages the multicast capabilities of NVSwitch to allow a GPU to send data to all other GPUs simultaneously with minimal communication steps. This results in only two synchronization steps being performed regardless of the number of GPUs involved, significantly increasing efficiency.

The process is divided into ReduceScatter tasks and AllGather tasks. Each GPU accumulates part of the resulting tensor and then broadcasts the accumulated result to all other GPUs. This method reduces per-GPU bandwidth and improves overall throughput.

Implications for AI Performance

Introducing TensorRT-LLM MultiShot can achieve nearly 3x speedup over existing methods, especially useful for scenarios that require low latency and high parallelism. These advancements allow for reduced latency or increased throughput at a given latency, potentially enabling ultra-linear scaling using more GPUs.

NVIDIA emphasizes the importance of understanding workload bottlenecks to optimize performance. The company is working closely with developers and researchers to implement new optimizations, with the goal of continuously improving the performance of the platform.

Image source: Shutterstock


Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Posts

Ether Lee (ETH) tests major support for $ 4,453 after the highest rejection.

August 31, 2025

Bitcoin analysts bet on $ 200K after hints of Fed.

August 23, 2025

‘Self -transactions, dressed in capital layout’: The cryptocurrency financial craze divides the industry.

August 15, 2025
Add A Comment

Comments are closed.

Recent Posts

RLUSD Stablecoin is extended to Africa to supply power to the border between the border.

September 5, 2025

Bybit Establishes New B2B Unit To Drive Institutional Adoption Of Digital Assets

September 5, 2025

Lowkick Studio Launches $SHARDS Token On Top Tier Exchanges For WorldShards MMORPG

September 5, 2025

The cryptocurrency is falling when the tokens and stocks connected to Trump are under pressure.

September 5, 2025

Cango Inc. Reports Second Quarter 2025 Unaudited Financial Results

September 5, 2025

Coindesk July 2025 Report: Stablecoins and CBDC

September 5, 2025

NOWPayments To Participate In SiGMA Europe Rome 2025

September 4, 2025

Web3 Enabler Announces Blockchain Payments V3.1 At Northeast Dreamin In Boston

September 4, 2025

Is XRP The Dark Horse Of The Cryptocurrency World? Earn 652 XRP Daily Using Invro Mining’s Smart Contract

September 4, 2025

TRX Was Early, ETH Set The Standard, BNB Built The Scale- Now SYC Brings The Next Evolution

September 4, 2025

Sign Up And Receive $500 Bonus, Ushering In A New Era Of Compliant And Secure Crypto Investment

September 4, 2025

Crypto Flexs is a Professional Cryptocurrency News Platform. Here we will provide you only interesting content, which you will like very much. We’re dedicated to providing you the best of Cryptocurrency. We hope you enjoy our Cryptocurrency News as much as we enjoy offering them to you.

Contact Us : Partner(@)Cryptoflexs.com

Top Insights

RLUSD Stablecoin is extended to Africa to supply power to the border between the border.

September 5, 2025

Bybit Establishes New B2B Unit To Drive Institutional Adoption Of Digital Assets

September 5, 2025

Lowkick Studio Launches $SHARDS Token On Top Tier Exchanges For WorldShards MMORPG

September 5, 2025
Most Popular

Trader predicts Altcoin Market’s relief rally, and one layer-1 crypto goes up and self.

February 24, 2025

Rock’s Not Dead in the Rockstar World Tour Hold and Win slot!

January 29, 2024

Rubin’s push for Rubens and CTV

February 8, 2024
  • Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms and Conditions
© 2025 Crypto Flexs

Type above and press Enter to search. Press Esc to cancel.