Post Content
In this video, we break down Poolside’s Laguna S2.1, an open-weights 118B MoE coding model (8B active per token) with a 1M-token context window, and why it performs above its size on agentic coding benchmarks like Terminal Bench 2.1. I cover how it was trained with reinforcement learning in FP8 on ~4,000 NVIDIA H200s in under nine weeks, plus how they tackled reward hacking on SWE-bench using an external LLM judge, prompt amendments, and network-blocked sandboxes.
Thanks to @NVIDIADeveloper for DGX Spark.
Laguna: https://poolside.ai/blog/introducing-laguna-s-2-1
DGX Spark: https://nvda.ws/3XIkwsh
Try it out: https://chat.poolside.ai/
Pool Agent Harness: https://poolside.ai/get-started
vLLM Serving: https://github.com/MiaAI-Lab/Laguna-S-2.1-DGX-Spark-RTX-6000-PRO
DSpark video: https://youtu.be/eFgknPFK-g0
MoE Quantization paper: https://arxiv.org/pdf/2606.00206
My voice to text App: whryte.com
Website: https://engineerprompt.ai/
RAG Beyond Basics Course:
https://prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
Let’s Connect:
🦾 Discord: https://discord.com/invite/t4eYQRUcXB
☕ Buy me a Coffee: https://ko-fi.com/promptengineering
|🔴 Patreon: https://www.patreon.com/PromptEngineering
💼Consulting: https://calendly.com/engineerprompt/consulting-call
📧 Business Contact: engineerprompt@gmail.com
Become Member: http://tinyurl.com/y5h28s6h
💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off).
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
00:00 Laguna S2.1 Overview
01:17 Benchmarks and Harness
02:13 Reinforcement Learning
03:34 Reward Hacking Fixes
05:16 Running on DGX Spark
06:55 NVFP4 Quantization
08:03 Speculative Decoding Speed
10:09 Pool Harness Demo
12:07 Verbose Reasoning Loops Read More Prompt Engineering
#AI #promptengineering