Google Speech GDX Plus

pub package MIT license GitHub issues

A community-maintained Flutter client for the Google Cloud Speech-to-Text API. It uses gRPC from Dart and supports synchronous recognition, streaming recognition, long-running recognition, API V1, V1p1beta1, V2, and experimental endless streaming.

Recognition demo

Recognition demo

Streaming demo

Streaming recognition demo

Requirements and compatibility

  • Dart 3.8 or later (below Dart 4)
  • Flutter 3.32 or later
  • Android, iOS, Linux, macOS, or Windows
  • A Google Cloud project with the Speech-to-Text API enabled

The package uses dart:io and does not support Flutter Web. Google Cloud availability, quotas, supported audio formats, and pricing are documented in the official Speech-to-Text documentation.

Installation

Add the package to your application:

flutter pub add google_speech_gdx_plus

Or add it to pubspec.yaml:

dependencies:
  google_speech_gdx_plus: ^6.0.0

Then import the public library:

import 'package:google_speech_gdx_plus/google_speech.dart';

Google Cloud setup

Create or select a Google Cloud project, enable the Speech-to-Text API, and configure credentials by following Google's client-library quickstart.

Important

Do not ship service-account private keys in a mobile or desktop application. For production client applications, obtain short-lived access credentials from a trusted backend and use ThirdPartyAuthenticator or SpeechToText.viaToken. Service-account loading is most appropriate for trusted environments and local development.

Authentication

Third-party authenticator

Use a trusted service to issue or refresh short-lived credentials:

final speechToText = SpeechToText.viaThirdPartyAuthenticator(
  ThirdPartyAuthenticator(
    obtainCredentialsFromThirdParty: () async {
      final json = await requestCredentialsFromMyBackend();
      return AccessCredentials.fromJson(json);
    },
  ),
);

Access token

You are responsible for refreshing the token before it expires:

final speechToText = SpeechToText.viaToken(
  'Bearer',
  '<access-token>',
);

Service account

Load credentials from a JSON file in a trusted environment:

import 'dart:io';
import 'package:google_speech_gdx_plus/google_speech.dart';

final serviceAccount = ServiceAccount.fromFile(
  File('path/to/service-account.json'),
);
final speechToText = SpeechToText.viaServiceAccount(serviceAccount);

You can also load the JSON string directly:

final serviceAccount = ServiceAccount.fromString(
  await rootBundle.loadString('assets/service-account.json'),
);
final speechToText = SpeechToText.viaServiceAccount(serviceAccount);

Never commit credential files. Bundling a service-account key as a Flutter asset exposes it to application users.

Recognize an audio file

Create a recognition configuration:

final config = RecognitionConfig(
  encoding: AudioEncoding.LINEAR16,
  model: RecognitionModel.basic,
  enableAutomaticPunctuation: true,
  sampleRateHertz: 16000,
  languageCode: 'en-US',
);

Load audio and send the request:

final audio = await File('test.wav').readAsBytes();
final response = await speechToText.recognize(config, audio);

for (final result in response.results) {
  print(result.alternatives.first.transcript);
}

Stream audio

Pass a Stream<List<int>> from a file or microphone source:

final streamingConfig = StreamingRecognitionConfig(
  config: config,
  interimResults: true,
);

final audioStream = File('test.wav').openRead();
final responseStream = speechToText.streamingRecognize(
  streamingConfig,
  audioStream,
);

responseStream.listen((response) {
  for (final result in response.results) {
    print(result.alternatives.first.transcript);
  }
});

API variants

  • Use SpeechToText for the stable V1 API.
  • Use SpeechToTextBeta for V1p1beta1 features.
  • Use SpeechToTextV2 for the V2 API. See example/audio_file_example_v2.
  • Use EndlessStreamingService, EndlessStreamingServiceBeta, or EndlessStreamingServiceV2 for experimental endless streaming.
final service = EndlessStreamingService.viaServiceAccount(serviceAccount);

service.endlessStream.listen((response) {
  // Handle recognition results.
});

service.endlessStreamingRecognize(
  StreamingRecognitionConfig(config: config, interimResults: true),
  audioStream,
);

Troubleshooting empty responses

An empty response is commonly caused by audio encoding, sample rate, channel, or language settings that do not match the supplied audio. Review Google's empty-response troubleshooting guide. Historical discussion is preserved in upstream issue #25.

Examples

Runnable examples are available in the example directory for file recognition, V2 recognition, microphone streaming, flutter_sound, and endless streaming. The examples intentionally reference the package through a local path when run from this repository.

Maintained Package

google_speech_gdx_plus is a community-maintained continuation of Felix Junghans's original google_speech project. This fork continues maintenance because the upstream package is inactive. Original copyright, license, commit history, and contributor credit remain intact.

Feature Requests

Feature requests and Pull Requests are always welcome.

Need Help?

Professional help is available for Google Cloud Speech integration, Flutter development, package and plugin maintenance, upgrades, and production troubleshooting. Visit gurwinderdevx.com to discuss a project.

Maintainer

Maintained by Gurwinder Singh.

Acknowledgements

The project was originally created and maintained by Felix Junghans. Thanks to Felix and all upstream contributors whose work established and improved the package.

License

Licensed under the MIT License. The original copyright notice is preserved.

Libraries

auth/third_party_authenticator
config/longrunning_result
config/recognition_config
config/recognition_config_v1
config/recognition_config_v1p1beta1
config/recognition_config_v2
config/streaming_recognition_config
endless_streaming_service
endless_streaming_service_beta
endless_streaming_service_v2
exception
generated/google/api/label.pb
generated/google/api/label.pbenum
generated/google/api/label.pbjson
generated/google/api/launch_stage.pb
generated/google/api/launch_stage.pbenum
generated/google/api/launch_stage.pbjson
generated/google/api/monitored_resource.pb
generated/google/api/monitored_resource.pbenum
generated/google/api/monitored_resource.pbjson
generated/google/cloud/speech/v1/cloud_speech.pb
generated/google/cloud/speech/v1/cloud_speech.pbenum
generated/google/cloud/speech/v1/cloud_speech.pbgrpc
generated/google/cloud/speech/v1/cloud_speech.pbjson
generated/google/cloud/speech/v1/resource.pb
generated/google/cloud/speech/v1/resource.pbenum
generated/google/cloud/speech/v1/resource.pbjson
generated/google/cloud/speech/v1p1beta1/cloud_speech.pb
generated/google/cloud/speech/v1p1beta1/cloud_speech.pbenum
generated/google/cloud/speech/v1p1beta1/cloud_speech.pbgrpc
generated/google/cloud/speech/v1p1beta1/cloud_speech.pbjson
generated/google/cloud/speech/v1p1beta1/resource.pb
generated/google/cloud/speech/v1p1beta1/resource.pbenum
generated/google/cloud/speech/v1p1beta1/resource.pbjson
generated/google/cloud/speech/v2/cloud_speech.pb
generated/google/cloud/speech/v2/cloud_speech.pbenum
generated/google/cloud/speech/v2/cloud_speech.pbgrpc
generated/google/cloud/speech/v2/cloud_speech.pbjson
generated/google/longrunning/operations.pb
generated/google/longrunning/operations.pbenum
generated/google/longrunning/operations.pbgrpc
generated/google/longrunning/operations.pbjson
generated/google/protobuf/any.pb
generated/google/protobuf/any.pbenum
generated/google/protobuf/any.pbjson
generated/google/protobuf/duration.pb
generated/google/protobuf/duration.pbenum
generated/google/protobuf/duration.pbjson
generated/google/protobuf/empty.pb
generated/google/protobuf/empty.pbenum
generated/google/protobuf/empty.pbjson
generated/google/protobuf/field_mask.pb
generated/google/protobuf/field_mask.pbenum
generated/google/protobuf/field_mask.pbjson
generated/google/protobuf/struct.pb
generated/google/protobuf/struct.pbenum
generated/google/protobuf/struct.pbjson
generated/google/protobuf/timestamp.pb
generated/google/protobuf/timestamp.pbenum
generated/google/protobuf/timestamp.pbjson
generated/google/protobuf/wrappers.pb
generated/google/protobuf/wrappers.pbenum
generated/google/protobuf/wrappers.pbjson
generated/google/rpc/status.pb
generated/google/rpc/status.pbenum
generated/google/rpc/status.pbjson
google_speech
speech_client_authenticator
speech_to_text
speech_to_text_beta
speech_to_text_v2