Recently I released Piper - Neural TTS, an iOS and macOS app that brings modern neural voices to Apple devices.
This project started as a simple question:
Why is it still so hard to use high-quality offline neural TTS on Apple platforms?
Most modern text-to-speech systems rely on cloud APIs. They require network access, introduce latency, and raise privacy concerns. But powerful open-source solutions already exist — they just aren't easily usable on iOS.
So I decided to fix that.
The Technology Behind It
The app is built around Piper, a fast neural text-to-speech engine designed to run locally.
Piper is an open-source TTS system that generates speech using neural models exported to ONNX Runtime, enabling efficient inference even on modest hardware.
Key characteristics:
- Fully offline
- Neural voices with natural prosody
- Optimized for CPU inference
- Multiple languages and voices
Unlike traditional concatenative TTS engines, neural systems generate speech with more natural rhythm and pronunciation.
This makes them suitable for:
- accessibility tools
- screen readers
- audiobook readers
- AI assistants
- productivity apps
Why Build This for Apple Platforms?
Despite Piper being widely used in Linux tools, smart-home systems, and accessibility software, support on Apple platforms is still limited.
iOS and macOS introduce several challenges:
1. Sandboxed Environment
Models must be downloaded, stored, and managed inside the app sandbox.
2. Performance Constraints
Neural TTS models can be tens of megabytes, so memory and CPU usage must be carefully managed.
3. System Integration
To be truly useful, a TTS engine must integrate with system accessibility features.
The goal of Piper – Neural TTS is to bridge modern neural speech synthesis with Apple’s accessibility ecosystem.
The app installs neural voices that can be used with VoiceOver and other system speech features, bringing clearer and more natural speech to everyday workflows. ([App Store][2])
Key Features
System-wide Voice Integration Install neural voices that can be used by accessibility features like VoiceOver.
Offline Neural Speech All synthesis happens locally on your device.
Multilingual Voices A growing collection of languages and regional accents.
Simple Voice Management Download and manage voices directly from the app.
Another important design goal: privacy.
No text leaves your device.
Engineering Challenges
Building this required solving several interesting technical problems:
Running ONNX Runtime on Apple platforms
Neural voice models are distributed as .onnx files, so efficient inference requires integrating ONNX Runtime.
Bridging native code with Swift
The Piper engine itself is written in C/C++, so a Swift-friendly interface had to be built around it.
Packaging neural models
Voice models are large (often ~50–60MB each), so the app needs a robust download and management system.
Phonemization
Piper relies on eSpeak-NG for phoneme generation, which adds another layer of integration.
Why Offline Speech Matters
Cloud speech services are powerful, but they have limitations:
| Cloud TTS | Local Neural TTS |
|---|---|
| Requires internet | Works offline |
| Potential privacy concerns | Fully private |
| API costs | Free once installed |
| Latency | Instant |
For accessibility tools, offline capability is especially important.
Your screen reader should work anytime, anywhere.
Open Source
The project is also available on GitHub:
https://github.com/IhorShevchuk/piper-app
The goal is to make neural TTS more accessible to developers and users across Apple platforms.
Try It
If you’re interested in offline neural voices on iOS or macOS, try the app:
Piper - Neural TTS
And if you're a developer working with speech technology or accessibility tools, I'd love to hear your feedback.
Email: ihor@ihor-shevchuk.dev