Found 13 articles
Auto-Tune Your LLM Judge
Measure the evaluator's variance, then rebuild its prompt from zero.

Announcing ExploitHunter.app
An open-source security workbench that keeps scope, approvals, tools, and evidence in one project.

Don't Fear the Model Router
Route to the best model with confidence.

Fight Evils with Evals!
Benchmarks measure benchmarks. Your system needs its own measures.

It's Time for llm:// Connection Strings
Simplify Model & Provider Config with llm:// URLs

Stop Asking LLMs to Do Math
They are bad at it. Here is how to fix it.

Stop Building Flaky Agents: Use Workflows & Memory
Deterministic patterns for non-deterministic models.

Your AI Agent is Useless Without This
Why MCP is the USB-C of Artificial Intelligence.

Production AI is Terrifying (And How to Fix It)
If your agent doesn't have guardrails, you aren't ready for production.

Don't Marry Your Model
LLM Routing, so hot right now

The Last to Think
Human or Machine?

Replacing Myself with AI
How I learned to stop worrying and love the GPT.

When AI Fails & the Crashing Robotic Cars
Google's self-driving cars are in 2x as many accidents as human drivers
