dvLatentSemanticTokens function

List<String> dvLatentSemanticTokens(
  1. String text, {
  2. Set<String> stopWords = dvLatentSemanticStopWords,
})

The words of text, lower-cased, without stop words, with the common English endings taken off so "deploying" and "deploys" are one word.

Implementation

List<String> dvLatentSemanticTokens(
  String text, {
  Set<String> stopWords = dvLatentSemanticStopWords,
}) {
  final List<String> out = <String>[];
  for (final RegExpMatch match
      in RegExp(r'[a-z0-9]+').allMatches(text.toLowerCase())) {
    final String word = match.group(0)!;
    if (word.length < 2 || stopWords.contains(word)) continue;
    out.add(_stem(word));
  }
  return out;
}