Sparking Transformation: How GPUs Busted Through a Once-Impossible Analytics Job
07/09/2021
A data scientist, she was tasked to comb a 3+ terabyte dataset at the Internal Revenue Service for patterns that might help uncover fraud. But even when she let the job run all night on a large bank of CPU servers the data refused to line up.
She returned in the morning to find the job had failed, so she tried again. It failed again.
About that time, Nasheb Ismaily of Cloudera knocked on the door of Rahul Tikekar, manager of a technical team that supports data analysts at the IRS. The Cloudera solutions engineer asked if Tikekar's team had any uses for Cloudera Data Platform (CDP), implementing Apache Spark 3.0 software accelerated by GPUs.
I jumped at the opportunity, said Tikekar. We have NVIDIA graphics cards on standalone servers, but using Spark to run them on a distributed cluster had eluded us for a while, so this was perfect timing for us and Deb had the perfect use case, he said.
A Nerdy Knot Untied A quick test of the software immediately speeded up many parts of Tylor's work up to 5x with no code changes, but a few pieces still lagged.
Ismaily called in a team of data scientists at NVIDIA to examine the guts of the code. They quickly determined a few tasks with particularly gnarly data structures were still running on CPUs. They wrote code to handle those jobs and inserted it into Spark's software interface for RAPIDS, the open library for running data analytics on GPUs.
Tylor ran another test, and boom, it all went on the GPUs in a distributed Spark cluster and the speedup was remarkable - Deb's running the whole program on a four-node cluster right now, said Tikekar.
The Cloudera and NVIDIA integration will empower us to use data-driven insights to power mission-critical use cases, said Joe Ansaldi, technical branch chief of the research and applied analytics and statistics division at the IRS and Tikekar's boss.
We're currently implementing this integration, and already seeing over 20x speed improvements at half the cost for our data engineering and data science workflows, he added.
Spark 3.0 + GPUs = New Horizons The work promises several payoffs the IRS team is already exploring.
With a Spark cluster of GPU-powered servers, the group can accelerate all its current jobs and run others previously thought impractical. And those jobs can tackle big datasets the team has at its disposal.
Before Spark 3.0, this was not possible, but now we're upping the ante with GPUs and we can dream of solving problems that were once impossible, said Tikekar.
Charting a Course to AI The team plans to apply what it learned with its success in data preparation, the so-called extract/transform/load (ETL) work of data analytics. Its next big step is accelerating full-blown AI inference jobs.
The partnership with Cloudera and NVIDIA helped us harness GPUs in clusters. When such advances come along, it takes a while to realize their power and develop apps that can use them, so Deb is really charting a new course for us - she's definitely the hero of the story, Tikekar said.
Specifically, the team aims to provide this distributed Spark-GPU infrastructure to analysts. Together, they will build large deep learning neural networks to tackle natural language processing and other analytics jobs currently impossible on a single server.
Many Apps for Machine Learning It's the kind of transformation many enterprises are seeking today with machine learning.
My personal feeling is that machine learning brings an incredible potential to make things that were difficult to achieve possible, said Tikekar, a Ph.D. in computer science who spent a decade teaching at Southern Oregon University before joining the IRS more than 13 years ago.
For example, today we scan in forms and then apply optical character recognition to read pieces of them, but with AI we can do a much better job of reading forms and finding patterns that can help find ID theft or reduce waste - a lot of applications can benefit from AI in numerous ways, he added.
To learn more about accelerating Cloudera's CDP 7.1.6 with NVIDIA GPUs, watch a GTC talk (free to view with registration) from October 2020, when the two companies announced their partnership.
And view Cloudera's demo below of a 44x speed increase on a data science workload using NVIDIA GPUs and RAPIDS compared to CPUs.
LINK: | https://blogs.nvidia.com/blog/2021/09/07/cloudera-spark-irs-gpus/... |
See more stories from nvidia |
More from Nvidia
22/04/2024
Climate Tech Startups Integrate NVIDIA AI for Sustainability Applications
Whether they're monitoring miniscule insects or delivering insights from satellites in space, NVIDIA-accelerated startups are making every day Earth Day. S...
18/04/2024
Wide Open: NVIDIA Accelerates Inference on Meta Llama 3
NVIDIA today announced optimizations across all its platforms to accelerate Meta Llama 3, the latest generation of the large language model (LLM). The open mod...
18/04/2024
Up to No Good: No Rest for the Wicked' Early Access Launches on GeForce NOW
It's time to get a little wicked. Members can now stream No Rest for the Wicked from the cloud. It leads six new games joining the GeForce NOW library of m...
18/04/2024
NVIDIA Honors Partners of the Year in Europe, Middle East, Africa
NVIDIA today recognized 18 partners in Europe, the Middle East and Africa for their achievements and commitment to driving AI adoption. The recipients were hon...
17/04/2024
Seeing Beyond: Living Optics CEO Robin Wang on Democratizing Hyperspectral Imaging
Step into the realm of the unseen with Robin Wang, CEO of Living Optics. The sta...
17/04/2024
Moving Pictures: Transform Images Into 3D Scenes With NVIDIA Instant NeRF
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and which showcases new hardware, ...
16/04/2024
New NVIDIA RTX A400 and A1000 GPUs Enhance AI-Powered Design and Productivity Workflows
AI integration across design and productivity applications is becoming the new s...
16/04/2024
To Cut a Long Story Short: Video Editors Benefit From DaVinci Resolve's New AI Features Powered by RTX
Editor's note: This post is part of our In the NVIDIA Studio series, which c...
15/04/2024
AI Is Tech's Greatest Contribution to Social Elevation,' NVIDIA CEO Tells Oregon State Students
AI promises to bring the full benefits of the digital revolution to billions acr...
10/04/2024
The Building Blocks of AI: Decoding the Role and Significance of Foundation Models
Editor's note: This post is part of the AI Decoded series, which demystifies...
10/04/2024
Combating Corruption With Data: Cleanlab and Berkeley Research Group on Using AI-Powered Investigative Analytics
Talk about scrubbing data. Curtis Northcutt, cofounder and CEO of Cleanlab, and ...
09/04/2024
NVIDIA Joins $110 Million Partnership to Help Universities Teach AI Skills
The Biden Administration has announced a new $110 million AI partnership between Japan and the United States that includes an initiative to fund research throug...
09/04/2024
Broadcasting Breakthroughs: NVIDIA Holoscan for Media, Available Now, Transforms Live Media With Easy AI Integration
Whether delivering live sports programming, streaming services, network broadcas...
09/04/2024
Start Up Your Engines: NVIDIA and Google Cloud Collaborate to Accelerate AI Development
NVIDIA and Google Cloud have announced a new collaboration to help startups arou...
04/04/2024
NVIDIA Ranked by Fortune at No. 3 on 100 Best Companies to Work For' List
NVIDIA jumped to No. 3 on the latest list of America's 100 Best Companies to Work For by Fortune magazine and Great Place to Work. It's the company'...
04/04/2024
The Elder Scrolls Online' Joins GeForce NOW for Game's 10th Anniversary
Rain or shine, a new month means new games. GeForce NOW kicks off April with nearly 20 new games, seven of which are available to play this week. GFN Thursday ...
03/04/2024
A New Lens: Dotlumen CEO Cornel Amariei on Assistive Technology for the Visually Impaired
Dotlumen is illuminating a new technology to help people with visual impairments...
03/04/2024
Coming Up ACEs: Decoding the AI Technology That's Enhancing Games With Realistic Digital Humans
Editor's note: This post is part of the AI Decoded series, which demystifies...
28/03/2024
Greater Scope: Doctors Get Inside Look at Gut Health With AI-Powered Endoscopy
From humble beginnings as a university spinoff to an acquisition by the leading global medtech company in its field, Odin Vision has been on an accelerated jour...
28/03/2024
Get Cozy With Palia' on GeForce NOW
Ease into spring with the warm, cozy vibes of Palia, coming to the cloud this GFN Thursday. It's part of six new titles joining the GeForce NOW library of ...
27/03/2024
Software Developers Launch OpenUSD and Generative AI-Powered Product Configurators Built on NVIDIA Omniverse
From designing dream cars to customizing clothing, 3D product configurators are ...
27/03/2024
NVIDIA Hopper Leaps Ahead in Generative AI at MLPerf
It's official: NVIDIA delivered the world's fastest platform in industry-standard tests for inference on generative AI. In the latest MLPerf benchmarks...
27/03/2024
Viome's Guru Banavar Discusses AI for Personalized Health
In the latest episode of NVIDIA's AI Podcast, Viome Chief Technology Officer Guru Banavar spoke with host Noah Kravitz about how AI and RNA sequencing are r...
27/03/2024
Unlocking Peak Generations: TensorRT Accelerates AI on RTX PCs and Workstations
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and which showcases new hardware, ...
26/03/2024
Boom in AI-Enabled Medical Devices Transforms Healthcare
The future of healthcare is software-defined and AI-enabled. Around 700 FDA-cleared, AI-enabled medical devices are now on the market - more than 10x the number...
26/03/2024
Model Innovators: How Digital Twins Are Making Industries More Efficient
A manufacturing plant near Hsinchu, Taiwan's Silicon Valley, is among facilities worldwide boosting energy efficiency with AI-enabled digital twins. A virt...
26/03/2024
Into the Omniverse: Groundbreaking OpenUSD Advancements Put NVIDIA GTC Spotlight on Developers
Editor's note: This post is part of Into the Omniverse, a series focused on ...
25/03/2024
NVIDIA Blackwell and Automotive Industry Innovators Dazzle at NVIDIA GTC
Generative AI, in the data center and in the car, is making vehicle experiences safer and more enjoyable. The latest advancements in automotive technology were...
21/03/2024
AI's New Frontier: From Daydreams to Digital Deeds
Imagine a world where you can whisper your digital wishes into your device, and poof, it happens. That world may be coming sooner than you think. But if you...
21/03/2024
You Transformed the World,' NVIDIA CEO Tells Researchers Behind Landmark AI Paper
Of GTC's 900+ sessions, the most wildly popular was a conversation hosted by...
21/03/2024
Instant Latte: NVIDIA Gen AI Research Brews 3D Shapes in Under a Second
NVIDIA researchers have pumped a double shot of acceleration into their latest text-to-3D generative AI model, dubbed LATTE3D. Like a virtual 3D printer, LATTE...
21/03/2024
Here Be Dragons: Dragon's Dogma 2' Comes to GeForce NOW
Arise for a new adventure with Dragon's Dogma 2, leading two new titles joining the GeForce NOW library this week. Set Forth, Arisen Fulfill a forgotten de...
20/03/2024
AI Decoded From GTC: The Latest Developer Tools and Apps Accelerating AI on PC and Workstation
Editor's note: This post is part of the AI Decoded series, which demystifies...
19/03/2024
NVIDIA Celebrates Americas Partners Driving AI-Powered Transformation
NVIDIA recognized 14 partners in the Americas for their achievements in transforming businesses with AI, this week at GTC. The winners of the NVIDIA Partner Ne...
19/03/2024
Climate Pioneers: 3 Startups Harnessing NVIDIA's AI and Earth-2 Platforms
To help mitigate climate change - one of humanity's greatest challenges - researchers are turning to AI and sustainable computing to accelerate and operatio...
19/03/2024
Secure by Design: NVIDIA AIOps Partner Ecosystem Blends AI for Businesses
In today's complex business environments, IT teams face a constant flow of challenges, from simple issues like employee account lockouts to critical securit...
19/03/2024
Generation Sensation: New Generative AI and RTX Tools Boost Content Creation
Editor's note: This post is part of our In the NVIDIA Studio series, which celebrates featured artists, offers creative tips and tricks, and demonstrates ho...
19/03/2024
NVIDIA, Huang Win Top Honors in Innovation, Engineering
NVIDIA today was named the world's most innovative company by Fast Company magazine. The accolade comes on the heels of company founder and CEO Jensen Huan...
18/03/2024
NVIDIA Edify Unlocks 3D Generative AI, New Image Controls for Visual Content Providers
NVIDIA Edify, a multimodal architecture for visual generative AI, is entering a ...
18/03/2024
From Atoms to Supercomputers: NVIDIA, Partners Scale Quantum Computing
The latest advances in quantum computing include investigating molecules, deploying giant supercomputers and building the quantum workforce with a new academic ...
18/03/2024
New NVIDIA Storage Partner Validation Program Streamlines Enterprise AI Deployments
A sharp increase in generative AI deployments is driving business innovation for...
18/03/2024
NVIDIA Unveils Digital Blueprint for Building Next-Gen Data Centers
Designing, simulating and bringing up modern data centers is incredibly complex, involving multiple considerations like performance, energy efficiency and scala...
18/03/2024
Generative AI Developers Harness NVIDIA Technologies to Transform In-Vehicle Experiences
Cars of the future will be more than just modes of transportation; they'll b...
18/03/2024
All Eyes on AI: Automotive Tech on Full Display at GTC 2024
All eyes across the auto industry are on GTC - the global AI conference running in San Jose, Calif., and online through Thursday, March 21 - as the world's ...
18/03/2024
All Aboard: NVIDIA Scores 23 World Records for Route Optimization
With nearly two dozen world records to its name, NVIDIA cuOpt now holds the top spot for 100% of the largest routing benchmarks in the last three years. And thi...
18/03/2024
We Created a Processor for the Generative AI Era,' NVIDIA CEO Says
Generative AI promises to revolutionize every industry it touches - all that's been needed is the technology to meet the challenge. NVIDIA founder and CEO ...
14/03/2024
NVIDIA GTC 2024: A Glimpse Into the Future of AI With Jensen Huang
NVIDIA's GTC 2024 AI conference will set the stage for another leap forward in AI. At the heart of this highly anticipated event: the opening keynote by Je...
14/03/2024
Reach for the Stars: Eight Out-of-This-World Games Join the Cloud
The stars align this GFN Thursday as more top titles from Ubisoft and Square Enix join the cloud. Star Wars Outlaws will be coming to the GeForce NOW library a...
13/03/2024
Currents of Change: ITIF President Daniel Castro on Energy-Efficient AI and Climate Change
AI-driven change is in the air, as are concerns about the technology's envir...
13/03/2024
AI Decoded: Demystifying Large Language Models, the Brains Behind Chatbots
Editor's note: This post is part of our AI Decoded series, which aims to demystify AI by making the technology more accessible, while showcasing new hardwar...