You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This derivative is available only for research and non-commercial use under the included BreezeBlue agreement. Gating records an acknowledgement but does not replace the agreement or supply voice and recording rights.

Log in or Sign Up to review the conditions and access this model content.

Instavar SG Narration Full SFT

Derived from Breeze TTS 2 by BreezeBlue and licensed for research and non-commercial use only.

This is an independent research checkpoint trained by Instavar on a consented, single-speaker subset of Singapore's National Speech Corpus. It is not an official BreezeBlue release and is not endorsed by BreezeBlue or RESONIA, INC.

Artifact

  • Base: BreezeBlue/Breeze-TTS-2
  • Pinned base revision: 799624c0b4a1daa8db6d28bbd9850043c0270734
  • Selected checkpoint: step 750 of 1,000
  • Updated synthesis parameters: 2,387,151,872
  • Frozen during training: text encoder and audio codec
  • Optimizer: FP32-master SGD

This repository contains inference roles only. Optimizer, scheduler, random state, trainer state, source recordings, caches, and private receipts are not included.

Use

Use the companion source toolkit:

python infer.py /models/sg-narration-full-sft \
  --text "The train arrives in five minutes." \
  --output output.wav

Source and documentation: instavar/breeze-tts2-finetuning

Experiment report: Breeze TTS 2 LoRA and Full Fine-Tuning for Singapore English

Related model: Instavar SG Narration LoRA R8

Collection: Breeze TTS 2 fine-tuning by Instavar

Evaluation

In a matched reference-free comparison, this checkpoint reached mean ECAPA speaker similarity 0.6973 and WER 0.0467. The LoRA adapter reached 0.6810 and the same WER. The paired ECAPA interval crossed zero, so the objective study did not establish a winner.

One listener preferred LoRA for cadence and long-form listening, with three short-prompt ties. Both releases mispronounced a tested local word. These results do not establish faithful identity cloning, general Singaporean-accent coverage, production fitness, or broad listener preference.

Licence and responsible use

The included BreezeBlue agreement permits research and non-commercial use only. Do not use this artifact for a product, paid service, client work, advertising, revenue generation, production deployment, impersonation, deception, or voice cloning without legally sufficient consent and recording rights.

Review LICENSE, NOTICE, PROVENANCE.json, and SHA256SUMS before use.

Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for instavar/sg-narration-full-sft

Finetuned
(7)
this model

Collection including instavar/sg-narration-full-sft

Evaluation results

  • Word error rate on FEMALE_01 matched reference-free prompts
    self-reported
    0.047
  • Mean ECAPA speaker similarity on FEMALE_01 matched reference-free prompts
    self-reported
    0.697