shisa.ai
Back to News
Product

Shisa AI Launches High-Accuracy, Ultra-Low-Latency Real-Time Speech Recognition API

A real-time speech recognition API combining high accuracy with ultra-low latency.

Shisa Inc. (Headquartered in Minato-ku, Tokyo; “Shisa AI”) has launched a real-time speech recognition API powered by its proprietary AI models.

The API combines high speech recognition accuracy with ultra-low latency, making it suitable for a wide range of applications where real-time response is essential, including live meeting transcription, call centers, and live streaming. All data processing is handled entirely on servers and infrastructure located in Japan, addressing the data security requirements of businesses and organizations.

Shisa AI’s real-time speech recognition API, “shisa-realtime-asr,” recorded a Character Error Rate (CER) of 5.0% on SPREDS, a Japanese speech recognition benchmark. The pay-as-you-go rate starts at US$0.45 per hour. With a monthly commitment of 2,000 hours, the rate is US$0.40 per hour; at 10,000 hours, it is US$0.36 per hour. For commitments of 50,000 hours or more, pricing is available upon request.

Engine processing latency is approximately 38.6 milliseconds. In more than 100 tests using diverse 10-second audio samples, the time from the end of speech to final result confirmation was approximately 254 milliseconds with a single connection and approximately 499 milliseconds with 10 concurrent connections. Accuracy and latency figures are reference values measured in Shisa AI’s test environment. Engine processing time and the time from the end of speech to final result confirmation are different metrics, and actual response times may vary depending on usage conditions.

All data processing, including the handling of audio data, is performed entirely on servers and infrastructure located in Japan. This helps reduce concerns around cross-border data transfers and provides an environment that is easier to evaluate for businesses and organizations with strict security requirements.

The API uses a clear and transparent pricing structure. See the pricing page for details.

Shisa AI has its own AI research and development team based in Japan. Feedback from users and businesses can be directly incorporated into model development, supporting continuous model updates that address Japanese-specific expressions, honorific language, regional dialects, and cultural nuances. The team can also provide technical support during implementation and respond quickly to customization requirements based on individual business needs.

The API supports not only standard Japanese but also regional dialects such as the Kansai dialect, providing speech recognition optimized for Japanese-language environments. Shisa AI plans to add support for additional regional dialects, including those spoken in Tohoku and Kyushu. The API also includes a Custom Vocabulary feature for registering company names, product names, people’s names, technical terms, and industry-specific expressions in advance.

The Shisa Real-time ASR API became available on September 1, 2026. Pricing starts at US$0.45 per hour, and supported languages are Japanese and English. Learn more on the service page, or contact contact@shisa.ai.

Shisa Inc. develops open models optimized for Japanese and English, as well as translation and speech APIs. The company focuses on Japanese-specific context, honorific language, and cultural nuances, and works to provide AI models and services that are practical for businesses and developers in Japan.

Official website: https://shisa.ai/ja/<br>Official X: https://x.com/shisa_ai<br>Official Facebook: https://www.facebook.com/Shisa.Inc<br>Official LinkedIn: https://jp.linkedin.com/company/shisa-ai

More from the newsroom