Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • ADOPTION
  • TRADING
  • HACKING
  • SLOT
  • CASINO
Crypto Flexs
  • DIRECTORY
  • CRYPTO
    • ETHEREUM
    • BITCOIN
    • ALTCOIN
  • BLOCKCHAIN
  • EXCHANGE
  • ADOPTION
  • TRADING
  • HACKING
  • SLOT
  • CASINO
Crypto Flexs
Home»ADOPTION NEWS»IBM Research Announces Innovations to Accelerate Enterprise AI Training
ADOPTION NEWS

IBM Research Announces Innovations to Accelerate Enterprise AI Training

By Crypto FlexsSeptember 23, 20243 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
IBM Research Announces Innovations to Accelerate Enterprise AI Training
Share
Facebook Twitter LinkedIn Pinterest Email

Jack Anderson
23 Sep 2024 03:32

IBM Research has introduced a new data processing technique that accelerates AI model training and significantly improves efficiency by leveraging CPU resources.





According to IBM Research, IBM Research has unveiled a groundbreaking innovation aimed at expanding the data processing pipeline for enterprise AI training. The advancement is designed to leverage the abundant capacity of CPUs to accelerate the creation of powerful AI models such as IBM’s Granite models.

Optimizing data preparation

Before training an AI model, a large amount of data needs to be prepared. This data often comes from various sources such as websites, PDFs, and news articles, and must go through several preprocessing steps. These steps include filtering out irrelevant HTML code, removing duplicates, and screening for abusive content. These tasks are important, but they are not limited by the availability of GPUs.

Petros Zerfos, principal research scientist for IBM Research’s Watsonx data engineering, emphasized the importance of efficient data processing. “A lot of the time and effort that goes into training these models is spent preparing the data for those models,” Zerfos said. His team has been drawing on expertise from a variety of domains, including natural language processing, distributed computing, and storage systems, to develop ways to improve the efficiency of the data processing pipeline.

CPU capacity utilization

Many steps in the data processing pipeline involve “embarrassingly parallel” computations, where each document can be processed independently. This parallelism allows the work to be distributed across multiple CPUs, which can significantly speed up data preparation. However, some steps, such as removing duplicate documents, require access to the entire data set, which cannot be done in parallel.

To accelerate IBM’s Granite model development, the team developed a process to rapidly provision and utilize tens of thousands of CPUs. This approach involved marshalling idle CPU capacity across IBM’s Cloud data center network to ensure high communication bandwidth between CPUs and data storage. Traditional object storage systems are often underperforming, leaving CPUs idle, so the team used IBM’s high-performance Storage Scale file system to efficiently cache active data.

AI Training Scaling

Last year, IBM scaled up to 100,000 vCPUs on IBM Cloud to process 14 petabytes of raw data, generating 40 trillion tokens for AI model training. The team automated these data pipelines using Kubeflow on IBM Cloud. Their method proved to be 24x faster than previous techniques with Common Crawl data processing.

All of IBM’s open source Granite code and language models are trained using data prepared through these optimized pipelines. IBM has also made a significant contribution to the AI ​​community by developing the Data Prep Kit, a toolkit hosted on GitHub that simplifies data preparation for large-scale language model applications, supporting pretraining, fine-tuning, and augmented search generation (RAG) use cases. Built on distributed processing frameworks such as Spark and Ray, the kit allows developers to build scalable custom modules.

For more information, visit the official IBM Research blog.

Image source: Shutterstock


Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Posts

‘Self -transactions, dressed in capital layout’: The cryptocurrency financial craze divides the industry.

August 15, 2025

As you challenge the mixed technology signal, OnDo Price Hovers challenges the August Bullish predictions.

August 7, 2025

XRP Open Interests decrease by $ 2.4B after recent sale

July 30, 2025
Add A Comment

Comments are closed.

Recent Posts

GEMINI has been disclosed by IPO, Tilecer Gemi’s NASDAQ listing plan

August 16, 2025

Ethereum-based Meme Coin Pepeto Nears Stage 10, Raises Over $6.18M In Presale, As Ethereum Eyes $10,000

August 15, 2025

Trump’s encryption reform pushes Bitcoin higher

August 15, 2025

Ether Leeum can increase to $ 15 million as the institution accumulates: Study

August 15, 2025

‘Self -transactions, dressed in capital layout’: The cryptocurrency financial craze divides the industry.

August 15, 2025

Mawari Partners With Caldera To Launch Mawari Network, Enabling Real-Time Streaming Of Immersive, AI-Powered Experiences Globally

August 15, 2025

Re -creation attack in ERC -1155 -Ackee Blockchain

August 14, 2025

QF Network Confirms Q4 2025 Mainnet Launch To Redefine Layer-1 Blockchain Performance

August 14, 2025

Bybit EU Taps XION For Inaugural Launchpool In The EU, Opening Regulated Access For 450M+ Users

August 14, 2025

XRP’s intersection: Global financial backbone or $ 190 billion fantasy?

August 14, 2025

Keepsolid launches KS COIN: Loyalty encryption through actual utility token benefits

August 14, 2025

Crypto Flexs is a Professional Cryptocurrency News Platform. Here we will provide you only interesting content, which you will like very much. We’re dedicated to providing you the best of Cryptocurrency. We hope you enjoy our Cryptocurrency News as much as we enjoy offering them to you.

Contact Us : Partner(@)Cryptoflexs.com

Top Insights

GEMINI has been disclosed by IPO, Tilecer Gemi’s NASDAQ listing plan

August 16, 2025

Ethereum-based Meme Coin Pepeto Nears Stage 10, Raises Over $6.18M In Presale, As Ethereum Eyes $10,000

August 15, 2025

Trump’s encryption reform pushes Bitcoin higher

August 15, 2025
Most Popular

Farmville Creator Defends Beta Release on Web3 Gaming

August 10, 2024

What is Spectrum (SPEC)? – Bitfinex Blog

May 3, 2024

Senator Lummis Drafts Bill to Allow US States to Hold Bitcoin

July 30, 2024
  • Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms and Conditions
© 2025 Crypto Flexs

Type above and press Enter to search. Press Esc to cancel.