Open-source phoneme-level English pronunciation assessment (Wav2Vec2 + DTW). Self-hosted alternative to Azure Pronunciation Assessment.
-
Updated
Aug 15, 2026 - Python
Open-source phoneme-level English pronunciation assessment (Wav2Vec2 + DTW). Self-hosted alternative to Azure Pronunciation Assessment.
基于 Qwen3-TTS、Faster-Whisper 与 PaddleOCR 的 AI 语音与视觉服务平台,为英语学习 APP 提供 TTS 文本转语音(含 WebSocket 流式合成)、ASR 语音识别、音素级发音评价(G2P + MFCC/DTW)与图片/文档 OCR 四大能力。基于 FastAPI 构建,统一 RESTful + WebSocket API,内置管理后台与 Gradio 调试界面,支持 ModelScope 模型下载与多档配置,最低 6GB 显存即可运行。
Official Python SDK for the Vocametrix voice analysis API — AVQI, DSI, jitter/shimmer, pronunciation assessment, speech-to-text, prosody similarity, and 40+ more clinical and acoustic measures.
AI-powered English speaking practice with real-time CEFR assessment. Whisper transcription, MFA pronunciation analysis, and a multi-model ensemble scoring the six CEFR competencies. FastAPI + Streamlit/React.
Official TypeScript / JavaScript SDK for the Vocametrix voice analysis API — AVQI, DSI, jitter/shimmer, pronunciation assessment, speech-to-text, prosody similarity. Works in Node and browsers.
Quran memorisation coach that grades spoken recitation word by word with a Whisper ensemble and adapts repetition to measured accuracy.
Official Model Context Protocol (MCP) server for Vocametrix — bring clinical voice analysis (AVQI, DSI, jitter/shimmer, pronunciation assessment, prosody similarity, and 40+ more) into Claude, Cursor, Zed, Windsurf, and any MCP-compatible agent.
Automatic speech evaluation toolkit for L2 pronunciation assessment
Phoneme-level pronunciation assessment using frozen Whisper embeddings + BiLSTM + Attention. Achieves 97% accuracy on 7-class classification with DTW-based pronunciation scoring.
PronounceAI - English Pronunciation Assessment Tool for Livo AI SWE Assessment. Upload 30-45s English audio and get instant feedback with score, transcript, and specific improvement suggestions. Built with Next.js 14, TypeScript & Tailwind. Due to API access limitations during development, realistic mock data is used to deliver a fully functional
한국어 학습자 발화의 발음 정확성·유창성(0~4점)을 평정하는 평정자 캘리브레이션 과제와 응답 수집 서버, 음성 전처리 스크립트. 코드만 공개하며 음성 파일은 포함하지 않습니다.
Offline pronunciation-assessment benchmark: forced alignment + boundary-aware GOP recover mispronunciations where naive scorers collapse under time-warp and coarticulation -- a 2x2 dissociation against known ground truth, numpy-only.
Local Sanskrit recitation coach for Bhagavad Gita shlokas with audio-based pronunciation analysis, shloka detection, practice mode, and LLM feedback.
한국어 난독증 읽기평가 엔진 - wav2vec2 음소 CTC + G2P 정렬 채점, 14종 임상 오류 분류, 발달 위계 기반 중재 처방 (목업 문항)
Add a description, image, and links to the pronunciation-assessment topic page so that developers can more easily learn about it.
To associate your repository with the pronunciation-assessment topic, visit your repo's landing page and select "manage topics."