GigaChat 3.5 Test Drive: Evaluating Model Capabilities for Business Tasks

GigaChat 3.5, released in July and made available with open weights in September, has undergone extensive testing. The authors conducted a comparative analysis of the model's performance in real-world business cases, using a proprietary benchmark to evaluate AI agent operations. The study compared GigaChat 3.5 against current Russian and Chinese alternatives to determine its suitability for enterprise applications. The 'AI agent races' provide insights into the competitiveness of this domestic model against global trends in large language models. The detailed findings regarding performance, accuracy, and efficiency in executing specific tasks are available in the full article.
This is a summary. Read the full article at the original source:
HabrRelated stories
700 manuscripts, 48 hours, three withdrawals. The verifier won.
OpenAI recently released a massive repository containing over 700 mathematical proofs generated by an internal AI model, promising to provide formaliz…
Co-op Legal Services faces backlash over AI employee surveillance
Co-op Legal Services has implemented an AI-driven monitoring system that records and evaluates every customer phone call made by its staff. Using Open…
Refugees in Kenya face dwindling pay and job instability as AI automates entry-level tech work
Refugees in Kenya’s Kakuma camp, who have long relied on digital microwork like data annotation and transcription, are facing a severe economic downtu…



