Perplexity AI leverages the NVIDIA inference stack to process 435 million queries per month.

Terrill Dickey
December 6, 2024 04:17

Perplexity AI leverages NVIDIA’s inference stack, including H100 Tensor Core GPUs and Triton Inference Server, to manage over 435 million search queries per month, optimizing performance and reducing costs.

Perplexity AI, a leading AI-powered search engine, successfully manages over 435 million searches every month thanks to NVIDIA’s advanced inference stack. According to NVIDIA’s official blog, the platform integrates NVIDIA H100 Tensor Core GPUs, Triton Inference Server, and TensorRT-LLM to efficiently deploy large language models (LLMs).

Provides multiple AI models

To meet diverse user needs, Perplexity AI operates more than 20 AI models simultaneously, including variants of the open source Llama 3.1 model. Each user request is matched to the best-fitting model using smaller classification models that determine user intent. These models are distributed across GPU pods, each managed by an NVIDIA Triton inference server, ensuring efficiency under strict service level agreements (SLAs).

Pods are hosted within a Kubernetes cluster with an internal frontend scheduler that directs traffic based on load and usage. This ensures consistent SLA compliance and optimizes performance and resource utilization.

Performance and cost optimization

Perplexity AI uses a comprehensive A/B testing strategy to define SLAs for a variety of use cases. This process aims to maximize GPU utilization while optimizing the cost of inference services while maintaining the target SLA. Smaller models focus on minimizing latency, while larger user-targeted models such as the Llama 8B, 70B, and 405B undergo detailed performance analysis to balance cost and user experience.

Performance is further improved by parallelizing model deployment across multiple GPUs and increasing tensor parallelism to lower servicing costs for latency-sensitive requests. This strategic approach allowed Perplexity to save approximately $1 million per year, exceeding the cost of third-party LLM API services, by hosting models on cloud-based NVIDIA GPUs.

Innovative technology for improved throughput

Perplexity AI is working with NVIDIA to implement ‘separate serving’, a method of separating inference stages to different GPUs to significantly increase throughput while complying with SLAs. This flexibility allows Perplexity to leverage a variety of NVIDIA GPU products to optimize performance and cost-effectiveness.

Further improvements are expected with the upcoming NVIDIA Blackwell platform, which promises significant performance gains through technological innovations including the second-generation Transformer Engine and advanced NVLink features.

Perplexity’s strategic use of the NVIDIA inference stack highlights the potential for AI-based platforms to efficiently manage massive query volumes and deliver high-quality user experiences while remaining cost-effective.

Image source: Shutterstock

Perplexity AI leverages the NVIDIA inference stack to process 435 million queries per month.

BNB holders gained 177% in 15 months through Binance Rewards Program.

ETH ETF loses $242M despite holding $2K in Ether

Hong Kong regulators have set a sustainable finance roadmap for 2026-2028.

How are cryptocurrency payments changing business cash flow and operations?

Cryptocurrency Inheritance Update: February 2026

Where ETH Holders Will Earn Daily Returns in 2026: Best Crypto Savings Accounts Review

Bybit Introduces Fixed-Rate UTA Loans Offering Up To 10x Leverage And Up To 180-Day Borrowing

Block Inc (XYZ) Adds 340 Bitcoin in Q4: Earnings Report

Intercepts $300M In Impersonalization, Scams And Frauds Via New AI-Driven Risk Framework

Bitcoin price recovery weakens and falls to $67,000 as prominent analyst predicts massive collapse.

Ethereum’s brutal price action contrasts with strong spot ETF demand. Will this spur a rebound?

AAVE Price Prediction: $137 Target by February 28 Amid Tech Recovery

A Free, Open-Source Validator Client With Built-In Acceleration For Solana

Best Crypto Presales Vs ICO Vs IDO – Complete 2026 Comparison Guide

Top Insights

How are cryptocurrency payments changing business cash flow and operations?

Cryptocurrency Inheritance Update: February 2026

Where ETH Holders Will Earn Daily Returns in 2026: Best Crypto Savings Accounts Review

Most Popular

Immutable (IMX) and YGG Power Web3 Gaming with $1M Player Rewards Initiative.

Bit Origin Secures $500 Million Equity And Debt Facilities To Launch Dogecoin Treasury

Sui selected as 2024 Blockchain Solution of the Year at AIBC Eurasia Awards

Perplexity AI leverages the NVIDIA inference stack to process 435 million queries per month.

Provides multiple AI models

Performance and cost optimization

Innovative technology for improved throughput

Related Posts