Technologies
Back
Artificial Intelligence & Machine Learning

How monday.com Runs Agent Evals Against Real Dependencies: Webinar Recap

Dev.to
Advertisement468 × 90
How monday.com Runs Agent Evals Against Real Dependencies: Webinar Recap

A recent webinar featuring Dor Cohen, AI Engineering Director at monday.com, explored the challenges of testing AI agents. The session highlighted why traditional mock-based testing often fails to capture real-world agent behavior, as mocks lack production-level data complexity and state management. To solve this, monday.com utilizes mirrord to connect their evaluation pipelines to live staging clusters, allowing agents to interact with real databases and services. This approach ensures that evaluations account for trajectory, tool precision, and actual goal completion rather than just final outputs. By integrating this infrastructure into local development, Slack-based agents, and CI/CD pipelines, the team can catch regressions—such as model performance drops—before they reach production. The discussion emphasizes that high-fidelity testing against real dependencies is essential for reliable agent deployment, effectively bridging the gap between simulated environments and the messy reality of production systems.

This is a summary. Read the full article at the original source:

Dev.to
Advertisement468 × 90
Share
Artificial Intelligence & Machine Learning

Related stories

Advertisement970 × 250