Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • TRADING
  • SUBMIT
Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • TRADING
  • SUBMIT
Crypto Flexs
Home»ADOPTION NEWS»NVIDIA launches Nemotron-CC, a large-scale dataset for LLM pre-training
ADOPTION NEWS

NVIDIA launches Nemotron-CC, a large-scale dataset for LLM pre-training

By Crypto FlexsJanuary 10, 20253 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
NVIDIA launches Nemotron-CC, a large-scale dataset for LLM pre-training
Share
Facebook Twitter LinkedIn Pinterest Email

Iris Coleman
January 10, 2025 14:13

NVIDIA launched Nemotron-CC, a 6.3 trillion token English dataset, powering pre-training of large-scale language models with innovative data curation methods.





NVIDIA announced the launch of Nemotron-CC, a groundbreaking 6.3 trillion token English language dataset designed to improve dictionary training of large-scale language models (LLMs). This dataset, derived from Common Crawl, aims to increase the accuracy and efficiency of LLM through innovative data curation techniques, including the use of 1.9 trillion synthetically generated data tokens, according to NVIDIA.

Enhancing LLM pre-education

NVIDIA’s initiative addresses a critical need in LLM training, where the quality of pre-training datasets plays a pivotal role. Recent models, such as Meta’s Llama series, have been based on datasets consisting of up to 15 trillion tokens, but the exact composition of these datasets is largely unknown. Nemotron-CC seeks to fill this gap by providing the wider community with high-quality datasets that can support both short-term and long-term Token Horizon training.

Existing datasets often sacrifice up to 90% of data to improve benchmark accuracy, limiting their usefulness for widespread training. However, Nemotron-CC demonstrates how advanced methods such as classifier ensembles and synthetic data reconstruction can transform the Common Crawl data into a superior dataset that outperforms the Llama 3.1 8B model.

important results

The efficacy of Nemotron-CC is demonstrated by its performance on a variety of benchmarks. When training an 8B parameter model on 1 trillion tokens, the high-quality subset Nemotron-CC-HQ outperforms key datasets such as DCLM, increasing the MMLU score by 5.6 points. Additionally, the full 6.3 trillion token dataset matches MMLU’s DCLM while providing 4x more unique real-world tokens. This allowed the Nemotron-CC trained model to outperform Llama 3.1 8B on several metrics, including a 5-point increase in MMLU and a 3.1-point increase in ARC-Challenge score, enabling effective training over long token periods.

Innovative data curation technology

The development of Nemotron-CC involved several key insights. By combining different model-based classifiers, NVIDIA was able to select a wider range of high-quality tokens. Paraphrasing techniques also reduces noise and errors, creating diverse and valuable data transformations. The decision to disable the existing heuristic filter further improved the quality of the data set without compromising accuracy.

NVIDIA leveraged the NeMo Curator tool to extract and refine data from Common Crawl, applying filters for language, deduplication, and quality classification. This process was complemented by synthetic data generation, contributing approximately 2 trillion tokens to the dataset.

future prospects

Nemotron-CC has established itself as an essential resource for pre-training cutting-edge LLMs across a diverse range of tokens. To further enhance LLM capabilities, NVIDIA plans to expand its product by releasing more specialized datasets, including datasets focused on specific areas such as mathematics.

Image source: Shutterstock


Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Posts

ETH Triple Top Rejects $2.4K as Analysts Show Weakness Against BTC

June 15, 2026

Google unveils Gemini Omni and Gemini 3.5 Flash AI models

May 30, 2026

These three Bitcoin charts say BTC price will recover to $82,000.

May 22, 2026
Add A Comment

Comments are closed.

Recent Posts

SEC specifies rules for tokenized securities

June 19, 2026

PremiumBlock Launches Non-Custodial Risk Hub For User-Created Prediction Markets, Perps And Web3 Poker

June 19, 2026

Ethereum Quantum-Proof Account Offer Could Make Wallet Protection Cheaper

June 19, 2026

Try to win on Great Game Rockies slots

June 18, 2026

Bitmine Immersion Technologies Announces Cash Dividend Of $0.1056 Per Share Of 9.50% Series A Perpetual Preferred Stock

June 18, 2026

Bitcoin Price Flashing Buy Signal: The Same Signal Is Being Delivered

June 18, 2026

Stratosphere, Pudgy Penguins And Streamex Host Founders Table VIP Dinner During ETHConf 2026 And NYC Tech Week

June 18, 2026

ORBS) Reports Total Holdings Of Approximately $472 Million, Includes OpenAI, Beast Industries, More Than 16,000 ETH And Over 283 Million WLD Tokens

June 18, 2026

Capital B shareholders have approved the ability to raise up to $120 billion in Bitcoin funding.

June 18, 2026

Calais Becomes 1st Quantitative Hedge Fund To Deploy UBS UMINT As OES Collateral Via Bybit, ByCustody & DigiFT

June 18, 2026

HBAR outperforms XLM and LINK Developing: Bullish Signal or Noise?

June 18, 2026

Crypto Flexs is a Professional Cryptocurrency News Platform. Here we will provide you only interesting content, which you will like very much. We’re dedicated to providing you the best of Cryptocurrency. We hope you enjoy our Cryptocurrency News as much as we enjoy offering them to you.

Contact Us : Partner(@)Cryptoflexs.com

Top Insights

SEC specifies rules for tokenized securities

June 19, 2026

PremiumBlock Launches Non-Custodial Risk Hub For User-Created Prediction Markets, Perps And Web3 Poker

June 19, 2026

Ethereum Quantum-Proof Account Offer Could Make Wallet Protection Cheaper

June 19, 2026
Most Popular

Google AI and robotic systems open new horizons in materials discovery

December 1, 2023

Anichess Surpasses 100,000 Monthly Players in Web3 Gaming Milestone

January 22, 2025

A make-or-break moment for Cardano is approaching.

December 28, 2023
  • Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms and Conditions
© 2026 Crypto Flexs

Type above and press Enter to search. Press Esc to cancel.