Agents Fail Twice: The Tool Breaks, Then They Report Success
FTA freezes failed tool observations before generation. Across six models and 3,600 human-annotated responses, false-success falls from 22.8% baseline to 0.8% with a structured evidence contract, while useful responses rise from 74.9% to 98.8%.