Post Content
Large models are powerful, but expensive to run in production. In this demo, we’ll show how teams use Foundry for distillation and supervised fine-tuning to train small language models for task‑specific accuracy, dramatically reducing latency and cost. We’ll cover when distillation makes sense, how it complements fine tuning and reinforcement learning, and what real production teams have learned when deploying smaller models at scale. Expect fast examples and lots of Q&A.
To learn more, please check out these resources:
* https://aka.ms/build/foundrydiscord
𝗦𝗽𝗲𝗮𝗸𝗲𝗿𝘀:
* William Liang
𝗦𝗲𝘀𝘀𝗶𝗼𝗻 𝗜𝗻𝗳𝗼𝗿𝗺𝗮𝘁𝗶𝗼𝗻:
This is one of many sessions from the Microsoft Build 2026 event. View even more sessions on-demand and learn about Microsoft Build at https://build.microsoft.com
DEM322 | English (US) | Working with models
Demo | (100) Foundational
#MSBuild
Chapters:
0:00 – Shift in AI goals from speed to cost-effective scalability
00:01:22 – Challenge: agents consume large token volumes
00:06:11 – Overview of Cleaning and Fine-Tuning Traces for Student Model
00:07:03 – Setup for Evaluation on 100 Tasks with Training and Holdout Split
00:11:44 – Fine-tuning improves student model performance and reasoning
00:12:45 – Transition to Foundry platform demonstration
00:17:08 – Case study: comparing base model and fine-tuned model behavior in refund scenario
00:19:16 – Scenario: Cancelling unshipped order and verifying refund policy logic
00:24:57 – Summary and gratitude—encouraging model distillation for cost-efficient operations Read More Microsoft Developer