Serverless GPU: Deploy your LLM seamlessly @ WAD25

Serverless GPU: Deploy your LLM seamlessly @ WAD25

Github source code: https://github.com/jlandure/simple-gemma3-ollama-langchainjs-app

Linkedin: https://www.linkedin.com/posts/jlandure_wwc25-gemma-googlecloud-activity-7348236927227097088-bZU3

Linkedin: https://www.linkedin.com/posts/jlandure_wad2025-serverlessgpu-wwc25-activity-7349324082124922882-Cj_R

AI solutions are booming, with best practices emerging, frameworks gaining popularity, and LLM switching becoming less cumbersome.

However, the path to production—especially when it comes to securely hosting your own LLM—often presents significant hurdles. This talk introduces an exciting solution: leveraging Serverless GPU options to minimize infrastructure overhead and maximize your focus on innovation.

We'll dive into how to use Google Cloud's Cloud Run with GPU to seamlessly deploy an open-source LLM, ensuring security and scalability without the usual headaches.

More Decks by Julien Landuré

See All by Julien Landuré

Kiro : Ne codez plus seul, pilotez vos agents de code @AWS Nantes

June 11, 2026 · AWS Nantes

Antigravity et l'agent manager, ou comment mixer Code & IA pour être vraiment productif @La Nuit des Communautés 2026

May 28, 2026 · Nuit des Communautés 2026

Récap Cloud Next 2026 @ GDG Cloud Nantes & GDG Rennes

May 5, 2026 · GDG Cloud Nantes & GDG Rennes