Field notes from building and operating AI systems in production: what breaks, why it happens, and how to test for it.
What breaks when open-weight models call tools in production, and how to test for it.
Qualifying models and building evals and tests you can trust.
When a provider retires your model: replacements, other hosts for the same weights, and routing.
General patterns for agent loops, tools and tiered models.
Cost tracking, alerting, monitoring and reliability for LLM features.
Validating, containing and caching what models return.
Retrieval and reranking for search and RAG pipelines.