Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • TRADING
  • SUBMIT
Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • TRADING
  • SUBMIT
Crypto Flexs
Home»ADOPTION NEWS»Inference Engine 2.0 with Together AI, Turbo, and Lite Endpoints Announced
ADOPTION NEWS

Inference Engine 2.0 with Together AI, Turbo, and Lite Endpoints Announced

By Crypto FlexsJuly 21, 20242 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
Inference Engine 2.0 with Together AI, Turbo, and Lite Endpoints Announced
Share
Facebook Twitter LinkedIn Pinterest Email

Terryl Dickey
18 July 2024 18:41

Together AI launches Inference Engine 2.0, offering Turbo and Lite endpoints for improved performance, quality, and cost efficiency.





Together AI announced the release of its new Inference Engine 2.0, which includes Turbo and Lite endpoints. This new inference stack is designed to deliver significantly faster decoding throughput and superior performance compared to existing solutions.

Performance Improvement

According to together.ai, Together Inference Engine 2.0 delivers 4x faster decoding throughput than open source vLLM and 1.3x to 2.5x faster than commercial solutions like Amazon Bedrock, Azure AI, Fireworks, and Octo AI. The engine achieves over 400 tokens per second on Meta Llama 3 8B thanks to advances in FlashAttention-3, faster GEMM and MHA kernels, quality-preserving quantization, and speculative decoding.

New Turbo and Lite endpoints

Together AI introduces new Turbo and Lite endpoints starting with Meta Llama 3. These endpoints balance performance, quality, and cost, allowing enterprises to avoid compromises. Together Turbo closely matches the quality of full-precision FP16 models, while Together Lite provides the most cost-effective and scalable Llama 3 models available.

Turbo endpoints provide fast FP8 performance while maintaining quality, are consistent with the FP16 reference model, and outperform other FP8 solutions on AlpacaEval 2.0. These Turbo endpoints are priced at $0.88 per million tokens for 70B and $0.18 per million for 8B, making them significantly cheaper than GPT-4o.

Together Lite endpoints provide high-quality AI models at a low cost using INT4 quantization, and for Llama 3 8B Lite, it is 6x cheaper than GPT-4o-mini at $0.10 per million tokens.

Adoption and Approval

More than 100,000 developers and companies, including Zomato, DuckDuckGo, and The Washington Post, are already leveraging the Together Inference Engine for their generative AI applications. Rinshul Chandra, COO of Food Delivery at Zomato, praised the engine for its high quality, speed, and accuracy.

Technological innovation

Together Inference Engine 2.0 incorporates several technological advancements, including FlashAttention-3, custom speculators, and quality-preserving quantization techniques. These innovations contribute to the engine’s superior performance and cost-effectiveness.

Future outlook

Together AI plans to continue pushing the boundaries of AI acceleration. The company aims to ensure that the Together Inference Engine remains at the forefront of AI technology by expanding support for new models, technologies, and kernels.

Turbo and Lite endpoints for the Llama 3 model are available starting today, with plans to expand to other models soon. Visit the Together AI pricing page for more information.

Image source: Shutterstock


Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Posts

Stellar (XLM) Highlights the Superiority of Native Tokenization in Securities

May 6, 2026

Bitcoin is at risk of liquidation of $1.4 billion if BTC rises to $80,000.

April 28, 2026

Polymarket Seeks $400 Million Raise to $15 Billion Valuation: Report

April 20, 2026
Add A Comment

Comments are closed.

Recent Posts

Real Assets Meet Digital Utility

May 12, 2026

Bitcoin Suisse Expands With Digital Asset License And Investment Business Act Registration Approval In Bermuda

May 12, 2026

Cantor8 Moves Deeper Into Africa’s Mobile Money Sector Via Yiksi Limited

May 12, 2026

Casper Network Publishes The Casper Manifest, A Multi-Year Roadmap To Power Regulated Real-World Assets And The Machine Economy

May 12, 2026

Bakkt switches to stablecoin infrastructure following 77% drop in Q1 revenue

May 12, 2026

$NXT Launches On OKX Boost, KuCoin, MEXC, And LBank — Bringing AI-Powered Global Entertainment To Web3

May 12, 2026

MEXC Launches Race To Zero Season 2 With A 2,000g Gold Bar Prize Pool

May 12, 2026

MultiBank Group’s Crypto Arm Mb.io Brings Ghana Gold On-chain With Kings Orbis, EON3 & Mavryk

May 11, 2026

Bitmine Immersion Technologies (BMNR) Announces ETH Holdings Reach 5.21 Million Tokens, And Total Crypto And Total Cash Holdings Of $13.4 Billion

May 11, 2026

Real-World Asset Tokenization: The Next Big Crypto Narrative?

May 11, 2026

Binance’s XRP whale retail spreads have fallen to 2024 levels. What’s going on?

May 10, 2026

Crypto Flexs is a Professional Cryptocurrency News Platform. Here we will provide you only interesting content, which you will like very much. We’re dedicated to providing you the best of Cryptocurrency. We hope you enjoy our Cryptocurrency News as much as we enjoy offering them to you.

Contact Us : Partner(@)Cryptoflexs.com

Top Insights

Real Assets Meet Digital Utility

May 12, 2026

Bitcoin Suisse Expands With Digital Asset License And Investment Business Act Registration Approval In Bermuda

May 12, 2026

Cantor8 Moves Deeper Into Africa’s Mobile Money Sector Via Yiksi Limited

May 12, 2026
Most Popular

As the leverage ratio rises, the price of Ethereum stagnates below $3,500. What are the next steps?

January 24, 2025

Bitcoin Cash, Solana, and DeeStream, which are trending upward, will likely see continued gains.

February 15, 2024

Crypto analyst explains how XRP could see a massive 4500% rise to $27.

January 28, 2024
  • Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms and Conditions
© 2026 Crypto Flexs

Type above and press Enter to search. Press Esc to cancel.