Post Content
Your AI coding assistant calls a 70B+ cloud model just to add a docstring. Snapdragon X2 Elite’s 80 TOPS NPU changes that. In this session, build a three-tier inference routing architecture—on‑device (≤13B), on‑prem (14B–34B), and cloud (70B+)—cutting cloud tokens by 67%, latency by 70%, and keeping most code local. Includes routing logic, quantization trade‑offs, and a deployable classifier.
Seating for this session is first-come, first-served. Add it to your schedule to plan your day and arrive early to secure a spot.
𝗦𝗽𝗲𝗮𝗸𝗲𝗿𝘀:
* Alberto Martinez
𝗦𝗲𝘀𝘀𝗶𝗼𝗻 𝗜𝗻𝗳𝗼𝗿𝗺𝗮𝘁𝗶𝗼𝗻:
This is one of many sessions from the Microsoft Build 2026 event. View even more sessions on-demand and learn about Microsoft Build at https://build.microsoft.com
BRKSP90 | English (US) | Developer tools & frameworks
Breakout | (200) Intermediate
#MSBuild
Chapters:
0:00 – Discussion on room energy and interactive Q&A setup
00:06:27 – Quantitative breakdown of token usage in coding tasks
00:09:43 – Analysis of workload complexity distribution using Claude Sonnet 4.6
00:15:26 – Cost savings potential from task division—up to $24K daily
00:22:06 – Illustrating 4x savings and economic justification challenges
00:25:00 – Introduction to tiered architecture and model complexity
00:31:21 – Research acknowledgments and final conclusion on 73% resolvable efficiency
00:33:27 – Optimization Framework: Measure Token Cost, Latency, and Iterate
00:37:12 – Vasion Question: Explaining the Fifth ‘Fallback’ Classifier Mechanism Read More Microsoft Developer