logo

NJP

Making AI Cheaper and Faster: Efficient Model Distillation

ServiceNow · Apr 01, 2026 · video

Welcome to the AI research bites. This series of short and informative talks showcases cutting-edge research work from ServiceNow AI Research team. The AI Research Bites are open to all, especially those interested in keeping up with the fast-paced AI research community.    Large Language Models are powerful but expensive and slow to run, especially as conversations get longer. In this presentation, Raymond Li explores how model distillation combined with new efficient architectures (like Mamba and DeltaNet) can deliver performant models at a fraction of the cost and latency, thus making AI more accessible and scalable. Paper: https://arxiv.org/pdf/2511.02651v1 ServiceNow AI Research team: https://www.servicenow.com/research/

View original source

https://www.youtube.com/watch?v=lgIyxR7HBEc