Robotics evals should test hardware, not just software
Article URL: https://techstatecraft.substack.com/p/robotics-evals-should-test-hardware Comments URL: https://news.ycombinator.com/item?id=49935731 Points: 1 #…
Conventional unit tests remain essential around an LLM application, but they do not answer the most important production question about whether a change preserved useful behavior. A parser can still return valid JSON while the answer becomes less grounded. A tool call can satisfy its schema while selecting the wrong tool. A retrieval pipeline can return documents successfully while omitting the evidence required for the final answer. This gap exists because an LLM application is not a deterministic function whose correctness can always be represented as actual == expected. Empirical work on code generation has shown substantial output variation across repeated calls, including at temperature zero, so single-run assertions can mistake sampling noise for either success or failure. Regression testing therefore has to evaluate behavior statistically and semantically, not only execution paths.
Unit Tests Prove Contracts, Not Behavior
Discussion (0)
No comments yet. Start the conversation!