Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • TRADING
  • SUBMIT
Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • TRADING
  • SUBMIT
Crypto Flexs
Home»ADOPTION NEWS»Perplexity AI leverages the NVIDIA inference stack to process 435 million queries per month.
ADOPTION NEWS

Perplexity AI leverages the NVIDIA inference stack to process 435 million queries per month.

By Crypto FlexsDecember 6, 20243 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
Perplexity AI leverages the NVIDIA inference stack to process 435 million queries per month.
Share
Facebook Twitter LinkedIn Pinterest Email

Terrill Dickey
December 6, 2024 04:17

Perplexity AI leverages NVIDIA’s inference stack, including H100 Tensor Core GPUs and Triton Inference Server, to manage over 435 million search queries per month, optimizing performance and reducing costs.





Perplexity AI, a leading AI-powered search engine, successfully manages over 435 million searches every month thanks to NVIDIA’s advanced inference stack. According to NVIDIA’s official blog, the platform integrates NVIDIA H100 Tensor Core GPUs, Triton Inference Server, and TensorRT-LLM to efficiently deploy large language models (LLMs).

Provides multiple AI models

To meet diverse user needs, Perplexity AI operates more than 20 AI models simultaneously, including variants of the open source Llama 3.1 model. Each user request is matched to the best-fitting model using smaller classification models that determine user intent. These models are distributed across GPU pods, each managed by an NVIDIA Triton inference server, ensuring efficiency under strict service level agreements (SLAs).

Pods are hosted within a Kubernetes cluster with an internal frontend scheduler that directs traffic based on load and usage. This ensures consistent SLA compliance and optimizes performance and resource utilization.

Performance and cost optimization

Perplexity AI uses a comprehensive A/B testing strategy to define SLAs for a variety of use cases. This process aims to maximize GPU utilization while optimizing the cost of inference services while maintaining the target SLA. Smaller models focus on minimizing latency, while larger user-targeted models such as the Llama 8B, 70B, and 405B undergo detailed performance analysis to balance cost and user experience.

Performance is further improved by parallelizing model deployment across multiple GPUs and increasing tensor parallelism to lower servicing costs for latency-sensitive requests. This strategic approach allowed Perplexity to save approximately $1 million per year, exceeding the cost of third-party LLM API services, by hosting models on cloud-based NVIDIA GPUs.

Innovative technology for improved throughput

Perplexity AI is working with NVIDIA to implement ‘separate serving’, a method of separating inference stages to different GPUs to significantly increase throughput while complying with SLAs. This flexibility allows Perplexity to leverage a variety of NVIDIA GPU products to optimize performance and cost-effectiveness.

Further improvements are expected with the upcoming NVIDIA Blackwell platform, which promises significant performance gains through technological innovations including the second-generation Transformer Engine and advanced NVLink features.

Perplexity’s strategic use of the NVIDIA inference stack highlights the potential for AI-based platforms to efficiently manage massive query volumes and deliver high-quality user experiences while remaining cost-effective.

Image source: Shutterstock


Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Posts

TRX Price Prediction: TRON targets $0.35-$0.62 despite the current oversold situation.

October 26, 2025

BTC RSI hits April low as Coinbase premium turns red.

October 18, 2025

Crypto Exchange Rollish is expanded to 20 by NY approved.

October 2, 2025
Add A Comment

Comments are closed.

Recent Posts

Open Miner Cloud Mining Revolutionizes Cryptocurrency Mining, Generating Up To $32,000 In Daily Profits.

October 31, 2025

Analysts predict a 1,500% rally when PEPE price reaches $0.00012.

October 30, 2025

Unibase (UB), Humanity (H), And ConstructKoin (CTK) Are This Week’s Crypto Winners As Decentralized Infra Shines

October 30, 2025

Let AI Work For You — Empowering Everyone To Profit From The Intelligence Era

October 30, 2025

NOWPayments Launches $0 USDT (TRC20) Network Fee Offer For New Partners

October 30, 2025

Jiuzi Holdings Launches $1 Billion Bitcoin Treasury With SOLV To Drive Institutional Yields And RWA Innovation

October 30, 2025

Hetu 3.0 – Deep Intelligence Money

October 30, 2025

Doodles has joined Universal Monsters and dropped a TON of NFT stickers.

October 30, 2025

Ethereum whales doubled down on ETH as the $5,000 price target moves higher.

October 30, 2025

SOL remains fixed below $200 despite surge in ETF trading volume

October 30, 2025

Bybit’s BbSOL Gains Institutional Custody Support From Anchorage Digital, Reinforcing Its Institutional-Grade Standing

October 30, 2025

Crypto Flexs is a Professional Cryptocurrency News Platform. Here we will provide you only interesting content, which you will like very much. We’re dedicated to providing you the best of Cryptocurrency. We hope you enjoy our Cryptocurrency News as much as we enjoy offering them to you.

Contact Us : Partner(@)Cryptoflexs.com

Top Insights

Open Miner Cloud Mining Revolutionizes Cryptocurrency Mining, Generating Up To $32,000 In Daily Profits.

October 31, 2025

Analysts predict a 1,500% rally when PEPE price reaches $0.00012.

October 30, 2025

Unibase (UB), Humanity (H), And ConstructKoin (CTK) Are This Week’s Crypto Winners As Decentralized Infra Shines

October 30, 2025
Most Popular

CME’s Bitcoin options open interest has reached an all-time high.

December 19, 2023

Fetch Price Prediction for Today, April 13 – FET Technical Analysis

April 14, 2024

Traders set up an Ether Leeum rival that can cause 2,915% rally.

March 17, 2025
  • Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms and Conditions
© 2025 Crypto Flexs

Type above and press Enter to search. Press Esc to cancel.