
Julien Salinas wears many hats. He's an entrepreneur, software developer and, until lately, a volunteer fireman in his mountain village an hour's drive from Grenoble, a tech hub in southeast France.
He's nurturing a two-year old startup, NLP Cloud, that's already profitable, employs about a dozen people and serves customers around the globe. It's one of many companies worldwide using NVIDIA software to deploy some of today's most complex and powerful AI models.
NLP Cloud is an AI-powered software service for text data. A major European airline uses it to summarize internet news for its employees. A small healthcare company employs it to parse patient requests for prescription refills. An online app uses it to let kids talk to their favorite cartoon characters.
Large Language Models Speak Volumes It's all part of the magic of natural language processing (NLP), a popular form of AI that's spawning some of the planet's biggest neural networks called large language models. Trained with huge datasets on powerful systems, LLMs can handle all sorts of jobs such as recognizing and generating text with amazing accuracy.
NLP Cloud uses about 25 LLMs today, the largest has 20 billion parameters, a key measure of the sophistication of a model. And now it's implementing BLOOM, an LLM with a whopping 176 billion parameters.
Running these massive models in production efficiently across multiple cloud services is hard work. That's why Salinas turns to NVIDIA Triton Inference Server.
High Throughput, Low Latency Very quickly the main challenge we faced was server costs, Salinas said, proud his self-funded startup has not taken any outside backing to date.
Triton turned out to be a great way to make full use of the GPUs at our disposal, he said.
For example, NVIDIA A100 Tensor Core GPUs can process as many as 10 requests at a time - twice the throughput of alternative software - thanks to FasterTransformer, a part of Triton that automates complex jobs like splitting up models across many GPUs.
FasterTransformer also helps NLP Cloud spread jobs that require more memory across multiple NVIDIA T4 GPUs while shaving the response time for the task.
Customers who demand the fastest response times can process 50 tokens - text elements like words or punctuation marks - in as little as half a second with Triton on an A100 GPU, about a third of the response time without Triton.
That's very cool, said Salinas, who's reviewed dozens of software tools on his personal blog.
Touring Triton's Users Around the globe, other startups and established giants are using Triton to get the most out of LLMs.
Microsoft's Translate service helped disaster workers understand Haitian Creole while responding to a 7.0 earthquake. It was one of many use cases for the service that got a 27x speedup using Triton to run inference on models with up to 5 billion parameters.
NLP provider Cohere was founded by one of the AI researchers who wrote the seminal paper that defined transformer models. It's getting up to 4x speedups on inference using Triton on its custom LLMs, so users of customer support chatbots, for example, get swift responses to their queries.
NLP Cloud and Cohere are among many members of the NVIDIA Inception program, which nurtures cutting-edge startups. Several other Inception startups also use Triton for AI inference on LLMs.
Tokyo-based rinna created chatbots used by millions in Japan, as well as tools to let developers build custom chatbots and AI-powered characters. Triton helped the company achieve inference latency of less than two seconds on GPUs.
In Tel Aviv, Tabnine runs a service that's automated up to 30% of the code written by a million developers globally (see a demo below). Its service runs multiple LLMs on A100 GPUs with Triton to handle more than 20 programming languages and 15 code editors.
https://blogs.nvidia.com/wp-content/uploads/2022/10/Tabnine.mp4
Twitter uses the LLM service of Writer, based in San Francisco. It ensures the social network's employees write in a voice that adheres to the company's style guide. Writer's service achieves a 3x lower latency and up to 4x greater throughput using Triton compared to prior software.
If you want to put a face to those words, Inception member Ex-human, just down the street from Writer, helps users create realistic avatars for games, chatbots and virtual reality applications. With Triton, it delivers response times of less than a second on an LLM with 6 billion parameters while reducing GPU memory consumption by a third.
It's another example of how LLMs are expanding AI's horizons.
Triton is widely used, in part, because its versatile. The software works with any style of inference and any AI framework - and it runs on CPUs as well as NVIDIA GPUs and other accelerators.
A Full-Stack Platform Back in France, NLP Cloud is now using other elements of the NVIDIA AI platform.
For inference on models running on a single GPU, it's adopting NVIDIA TensorRT software to minimize latency. We're getting blazing-fast performance with it, and latency is really going down, Salinas said.
The company also started training custom versions of LLMs to support more languages and enhance efficiency. For that work, it's adopting NVIDIA Nemo Megatron, an end-to-end framework for training and deploying LLMs with trillions of parameters.
The 35-year-old Salinas has the energy of a 20-something for coding and growing his business. He describes plans to build private infrastructure to complement the four public cloud services the startup uses, as well as to expand into LLMs that handle speech and text-to-image to address applications like semantic search.
I always loved coding, but being a good developer is not enough: You have to understand your customers
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
06/09/2026
June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...
19/08/2026
As the cost of IP nears price parity with traditional SDI, Reynolds argues that ...
19/08/2026
A private 5G bubble was chosen instead of traditional broadcast RF for the race ...
19/08/2026
(L-R) Rafael Manuel and Isabel Sicat attend the Filipi ana premiere during the 2026 Sundance Film Festival at The Ray Theatre on January 23, 2026, in Park Cit...
19/08/2026
Studio owners invited to have their say on rates cuts
The Music Producers Guild (MPG) are inviting UK-based studios to take part in an in-depth survey that ...
19/08/2026
Vintage-inspired desktop interface revealed
Harrison's latest hardware offering sees them head into the world of audio interfaces, delivering a 10-in/12...
19/08/2026
Granular effect remains locked to scale and tempo
Described as your dream granular effect , Minimal Audio's latest creation is capable of turning any s...
19/08/2026
First polyphonic SWAM instrument incoming
Audio Modeling's latest physical-modelling orchestral instrument is just around the corner, and is now availab...
19/08/2026
World Cup Games Generate 84 Billion Minutes Across FOX, Fox Sports 1 and NBCU...
19/08/2026
July brought typical peak-holiday drops in viewership, driven primarily by viewers stepping away from screens and a visible decline in reach across both traditi...
19/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
19/08/2026
Henry joins BCE as the company accelerates its commercial strategy and expands its international presence across the media, broadcast and technology markets. He...
19/08/2026
iWedia, a leading provider of software solutions for connected TV devices, today announced a major milestone in Android platform optimization with the successfu...
19/08/2026
NTP Technology announces that ST 2110-30 audio-over-IP connectivity with AES67 and RAVENNA capabilities is now available for the brand's world-class audio i...
19/08/2026
Net Insight has been selected to provide its Nimbra platform by the service provider Big Blue Marble, formerly ORS Group, as the core media transport solution f...
19/08/2026
LiveU has supported Coventry City FC and several other top-flight UK football clubs through one of their most demanding live production challenges of the year, ...
19/08/2026
Nature Seekers Documentary Shot with Blackmagic URSA Cine 12K LF
Brie Clayton August 18, 2026
0 Comments
Filming endangered turtles required camera...
19/08/2026
Neyrinck Brings D-Control Consoles Back to Life on Apple Silicon with D-Control ...
19/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
19/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
19/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
19/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
19/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
19/08/2026
Riedel Communications today announced that Edith Cowan University (ECU) in Perth, Australia, implemented a comprehensive communications infrastructure that brin...
19/08/2026
QuickLink, a leading provider of award-winning video production and remote guest contribution solutions, will spotlight a major leap forward in remote guest con...
19/08/2026
Integrated platform melds media asset management, playout automation and newsroom workflows
EMAM, PlayBox Neo and SNEWS have officially launched Nuvex, a conne...
19/08/2026
Screen Australia releases the Private Investment Toolkit to strengthen screen in...
19/08/2026
Do Aussies truly value local content? Unpacking Screen Currency 2026 with Deirdr...
19/08/2026
There's an early Christmas gift for Bookish fans, as UKTV confirms the re-commission of the hit U&alibi drama, created by and starring Mark Gatiss. Viewers ...
19/08/2026
Wednesday 19 August 2026
Sky gives sports fans the chance to watch multiple liv...
19/08/2026
On Saturday September 5 Katie Taylor will step into the ring for the fight of her life against Flora Pili in front of 80,000 fans in Croke Park. A legendary fig...
18/08/2026
SportStream 2026, a free three-day virtual event, will take place September 15-17. One registration provides access to all three days of live programming and on...
18/08/2026
BeckTV will introduce BeckFlow to the European market at IBC2026, demonstrated on the Providius stand (10.F53).
BeckFlow is a web-based schematic documentation...
18/08/2026
Boland Communications will introduce a SMPTE ST 2110 input option for its 4K monitors at IBC2026, along with a new 55-inch super-narrow bezel display for video ...
18/08/2026
Dolby Laboratories and the Seattle Seahawks have announced a partnership integrating Dolby OptiView into the Seahawks' live streaming coverage, making the S...
18/08/2026
For BMG Executive Producer of Sports Production, Graham Taylor, producing live s...
18/08/2026
TMRW Sports has appointed Rufus Hack as President of Golf, a newly created role in which he will lead the company's golf businesses including TGL presented ...
18/08/2026
SMT will provide race technology for FOX Sports' coverage of the Freedom 250 Grand Prix of Washington, D.C. on August 23, when the NTT INDYCAR SERIES races ...
18/08/2026
The live, YouTube-native roundtables combine fan voices, credentialed reporters,...
18/08/2026
Behind The Mic provides a roundup of recent news regarding on-air talent, includ...
18/08/2026
The industry legend sounds off on replacing the Heat's iconic Medusa score...
18/08/2026
William F. Rasmussen, the sports fan-turned broadcaster and entrepreneur whose out-of-the-box idea of a 24-hour cable sports network, ESPN, changed forever the ...
18/08/2026
DAZN Group has completed its acquisition of ViewLift, a provider of streaming and digital solutions for sports content owners. The acquisition was first announc...
18/08/2026
The American Association of Professional Baseball (AAPB) has added SI TV, Sports Illustrated's 24/7 streaming channel, as a broadcast partner for the remain...
18/08/2026
The Esports Foundation has announced that the inaugural Esports Nations Cup (ENC), originally scheduled for November 2026 in Riyadh, will be postponed to Novemb...
18/08/2026
Diversified has announced a new Global Managed Services Practice, consolidating the company's managed services capabilities into one organization. Rob Mello...
18/08/2026
Appear has completed a proof of concept with Thoroughbred Racing Productions (TRP), demonstrating multi-camera remote live production over bonded 5G cellular an...