跳转到内容

Developing Cadenza

此内容尚不支持你的语言。

Everything lives in cadenza/. The app is plain Swift (Swift 5 language mode, SwiftUI plus AppKit) compiled with swiftc; there is no Xcode project. Local speech models run through sherpa-onnx.

  • macOS 14 or later (Apple silicon or Intel), with the Xcode Command Line Tools (xcode-select --install). The build produces one binary for both chips (CADENZA_ARCHS=arm64 ./cadenza/build.sh --stage-only builds only the chip you are on, faster).
  • Optional: the sherpa-onnx static library for local models. Without it the app still builds, without local recognition.
Terminal window
./cadenza/tools/fetch-sherpa-onnx.sh # optional, once: downloads the inference library into cadenza/third_party
./cadenza/build.sh --stage-only # builds cadenza/build/stage.noindex/Cadenza.app.zip

--stage-only does not install anything and works without the maintainer’s signing certificate: the package is signed ad hoc. Plain ./cadenza/build.sh installs over /Applications/随言.app and is for maintainers; it refuses to run without the real signing identity so that a stray build can never replace the signature the installed app’s privacy permissions are tied to. CADENZA_SIGN_IDENTITY=<id> selects another certificate; - means ad hoc.

The staged app is a zip on purpose, so that no launchable copy sits in your checkout. To run it, unzip into a temporary directory, use it, then delete it:

Terminal window
T=$(mktemp -d) && unzip -q cadenza/build/stage.noindex/Cadenza.app.zip -d "$T"
"$T"/*.app/Contents/MacOS/Cadenza --selftest

The product is Cadenza (English) and 随言 (Simplified Chinese); user-visible names always come from Brand.name. The executable, Bundle ID (local.cadenza.app), Keychain service and support folder (~/Library/Application Support/Cadenza) all use the current name. Installs from before the rename keep working: LegacyMigration.swift moves the old folder, preferences and saved credentials once (macOS may ask once to allow access to the old Keychain items; choose Always Allow). The permissions the app needs (microphone, accessibility, input monitoring, screen recording) are granted again once, because macOS ties them to the Bundle ID. The code that recognises Keychain labels written under retired product names (Capsule.swift) also stays.

Run every suite before a pull request. Each prints failures=0 when it passes:

Terminal window
BIN="$T"/*.app/Contents/MacOS/Cadenza
"$BIN" --selftest # main suite
"$BIN" --selftest-settings-ui
"$BIN" --selftest-settings-polish
"$BIN" --selftest-status-menu
"$BIN" --selftest-asr-settings
"$BIN" --selftest-asr-entry
"$BIN" --selftest-local-model
"$BIN" --selftest-trigger-stage2
"$BIN" --check-brand-resources # both languages have the same keys
./cadenza/trigger-state-machine/run-tests.sh
  • Self-tests, previews, benchmarks and --check-* runs never ask macOS about permissions (src/TCC.swift): a build signed differently from the installed app (ad hoc) that does so makes macOS reset the installed app’s Accessibility, Input Monitoring and Screen Recording grants. Ask for permissions only through TCC; tools/check-tcc-calls.sh enforces it. Do not launch a staged build without one of those flags.
  • --selftest-models-ui opens the real settings window on a scratch configuration and an in-memory Keychain, and drives the “My AI models” controls (add, type a name, switch consent, save and delete a key, delete with confirmation) through the accessibility tree, as VoiceOver would. It needs a logged-in graphical session, so it is not in CI; run it by hand after changing that page. --selftest-sidebar-click does the same for the sidebar.
  • Self-tests never open windows and use isolated configuration. They do write lines into the local log file.
  • --selftest-local-model-real runs a real installed model on speech synthesized by macOS say. It skips itself when no model is installed. Download models inside the app (Settings → Speech → Local models).
  • Tests with a fake recorder or fake socket prove the protocol, not a real microphone or provider. Say which you ran.

--accuracy-benchmark speaks a fixed Mandarin corpus with several system voices under several conditions and scores each installed model (and, if you name them, cloud providers) with the character error rate. See ACCURACY.md for options and the current results. Use it to justify any change to recognition settings or any new model: keep a change only if it measurably helps.

Terminal window
"$BIN" --preview-brand-page=engines --preview-engine-tab=tuning \
--preview-brand-output=/tmp/engines.png --preview-brand-height=1100 --ui-language=zh-Hans

Pages: onboarding-1…, engines, engines-local, general, developer, capsule-record, capsule-error and more (BrandPreview.swift). Previews render off-screen with temporary configuration: no shortcut listeners, no recording, no network.

Path Contents
src/ All app source, plus *Fixtures.swift self-test suites
src/ASRCore.swift, ASRWire.swift, CloudASRRecorder.swift, *Provider.swift Cloud providers and their wire formats
src/LocalASR.swift, LocalModelCatalog.swift, LocalModelCenter.swift Local models: loading, catalog, download
src/VoicePipeline*.swift, HotkeyCenter.swift, Shortcut.swift Hold-to-talk pipeline and shortcut handling
src/MainSettingsView.swift, *View.swift SwiftUI settings and onboarding
src/LocalAPI*.swift Loopback developer API (LOCAL-API.md)
resources/ Localized strings, brand assets (resources/brand), illustrations
trigger-state-machine/ The trigger state machine, a separate Swift package with its own tests
docs/ Technical documentation
tools/ Helper scripts, including check-no-secrets.sh

The app currently ships English and Simplified Chinese, and the language list is not yet data-driven. A translation is more than a strings file: the list of languages is written out in code in several places. Today a third language needs:

  1. resources/<code>.lproj/Localizable.strings and InfoPlist.strings, translated from resources/en.lproj. Keep every key, keep %@/%d/%.1f placeholders in the same order, and keep product names out of the strings (they come from Brand.name).
  2. A new case in AppLanguage and its title key (src/Brand.swift), the language list inside L10n.verifyResources (it currently compares English with Simplified Chinese only), the bundled copies in build.sh, the localized privacy and terms documents (PRIVACY.<code>.md, TERMS.<code>.md, see PrivacySupport.swift), and the per-language names in LocalModelCatalog.swift.
  3. --check-brand-resources must pass, and the busiest pages should be previewed with --ui-language=<code>.

Making the language list data-driven, so that a translation is only a strings file plus one table entry, is on the roadmap and would be a very welcome contribution. If you want to translate, open an Add a language issue first so that we can coordinate this work.

Edit LocalModelCatalog.builtin and keep docs/models-manifest.example.json identical (a test compares them). A model needs a compatible license, download size, SHA-256 of every file, the files the loader requires and the languages it covers. Run --accuracy-benchmark against the recommended model and put the numbers in your pull request; also report speed, memory and whether it writes punctuation. The app never downloads a model without the user’s action.

A provider needs: a wire implementation in ASRWire.swift or its own file, option validation in ASRSettings.swift, credential fields in ASRCore.swift (stored in the Keychain), the configuration sheet, an upload-consent check, and wire-level tests with a fake socket covering consent, request construction, partial/final handling, cancellation and errors. Check the official documentation for every parameter you send. Mark in the pull request that the provider was tested with a fake transport unless you also ran it against a real account.

Match the surrounding code: naming, comment density and idioms. User-facing strings go through L10n.tr. Do not add dependencies without discussing them first; third-party code needs a license entry.