This is a autopost bolg frinds we are trying to all latest sports,news,all new update provide for you
Wednesday, September 11, 2024
Show HN: Tune LLaMa3.1 on Google Cloud TPUs https://ift.tt/Y1rKRxe
Show HN: Tune LLaMa3.1 on Google Cloud TPUs Hey HN, we wanted to share our repo where we fine-tuned Llama 3.1 on Google TPUs. We’re building AI infra to fine-tune and serve LLMs on non-NVIDIA GPUs (TPUs, Trainium, AMD GPUs). The problem: Right now, 90% of LLM workloads run on NVIDIA GPUs, but there are equally powerful and more cost-effective alternatives out there. For example, training and serving Llama 3.1 on Google TPUs is about 30% cheaper than NVIDIA GPUs. But developer tooling for non-NVIDIA chipsets is lacking. We felt this pain ourselves. We initially tried using PyTorch XLA to train Llama 3.1 on TPUs, but it was rough: xla integration with pytorch is clunky, missing libraries (bitsandbytes didn't work), and cryptic HuggingFace errors. We then took a different route and translated Llama 3.1 from PyTorch to JAX. Now, it’s running smoothly on TPUs! We still have challenges ahead, there is no good LoRA library in JAX, but this feels like the right path forward. Here's a demo ( https://ift.tt/XI6x4pf ) of our managed solution. Would love your thoughts on our repo and vision as we keep chugging along! https://ift.tt/hPQfkaG September 11, 2024 at 08:44PM
Subscribe to:
Post Comments (Atom)
Show HN: Inertia – Not your Grandfathers animation editor https://ift.tt/0PbBOYa
Show HN: Inertia – Not your Grandfathers animation editor https://ift.tt/K37HyMS August 9, 2026 at 03:57AM
-
Show HN: I built Dirac, Hash Anchored AST native coding agent, costs -64.8 pct Fully open source, a hard fork of cline. Full evals on the gi...
-
Show HN: Pixel text renderer using CSS linear-gradients (no JavaScript) I've been playing around with rendering pixel text using only CS...
-
Show HN: Total Recall – write-gated memory for Claude Code https://ift.tt/G7AugiK February 6, 2026 at 05:26AM
No comments:
Post a Comment