Documented LLM parameter timing for local model optimization
Welcome back
Documented the distinction between load-time and inference-time LLM parameters for local model serving.
Derived from this session's token and cost data. Not shown on the feed.
Comments