I Built Two Agent Systems. Each One Proved the Other One Wrong.

Developer Debashish Ghosal shares findings from building two distinct LLM-based agent systems, concluding that relying on prompts for safety is fundamentally flawed. In his 'PlannerCritic' and 'AdversarialDebate' projects, Ghosal initially attempted to solve behavioral issues through prompt engineering, such as instructing models to be 'adversarial' or 'independent.' Both approaches failed, resulting in unreliable verdicts and 'theatrical' outputs. Ghosal discovered that the solution for both systems was to move safety boundaries from the prompt layer into deterministic code. By enforcing structural invariants—such as using code-based severity contracts and mechanical isolation—he achieved significantly more reliable results. The author argues that because LLMs are inherently non-deterministic, developers should stop trying to make them reliable through prompts and instead build architectures where the model's output is irrelevant to the system's core safety boundaries.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
Mistral AI raises €3B led by Samsung: how sovereign is it?
Paris-based Mistral AI has secured a €3 billion Series D funding round, pushing its post-money valuation beyond €21 billion. Led by Samsung Electronic…
The author analyzes the current hype surrounding generative AI tools in software development. The text raises the question of the difference between q…
Engram is a sampler that turns broken AI hallucinations into music
Music startup Thoughtful Things has launched a Kickstarter campaign for Engram, a novel sampler and groovebox that integrates artificial intelligence…



