Shisa AI Releases New Benchmark “COSPA”
A new benchmark for measuring the cost of solving programming assignments.
On Friday, September 4, QwenMeetupTokyo was held, bringing together the latest technologies in generative AI and practical use cases centered around Alibaba’s large language model, Qwen.
At the event, Jia Shen, one of the co-founders of Shisa.AI, took the stage to share Shisa.AI’s current work on Speech Recognition, as well as the newly developed Cost Of Solving Programming Assignments (COSPA) benchmark.
COSPA is designed to make the cost of solving programming assignments measurable. There are already many rankings showing which models achieve the highest scores, but they have not shown how much it costs to reach those scores. Shisa.AI has built a benchmark that evaluates this cost as well.
COSPA uses more than 300 real-world agentic tasks, and running the full set of evaluations takes at least one full day in real time. With this benchmark, models can now be compared not only based on which model is the smartest, but also on how much cost and resources are required to achieve that level of performance.
You can learn more and try it out in the COSPA GitHub repository.
More from the newsroom
Shisa 7B released
A bilingual general-purpose chat model using a synthetic-data driven approach.
Read moreShisa-Gamma-7b-v1 Surpasses 1 Million Downloads
One year after its role in pioneering evolutionary model merges, our model reaches a significant milestone.
Read more