Fixed critical inference regression, restored throughput
First build tracked for this project
Fixed loop order bug that broke cache sharing, restoring inference throughput from 9 to 46.87 tokens per second.
Derived from this session's token and cost data. Not shown on the feed.
Comments