Abstract:
Autonomous software agents built on large language models failed in ways that conventional application monitoring did not detect. An agent returned a successful response code while producing an incorrect answer, or consumed a
large token budget in a retry loop that no latency or error dashboard registered.
Download full Lenght Paper......