Skip to content

Bring back character q-grams as a token generator - #52

Merged
sadit merged 1 commit into
mainfrom
qgram-generator
Sep 24, 2026
Merged

sadit merged 1 commit into
mainfrom
qgram-generator

Conversation

@sadit

@sadit sadit commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Summary

Adds QgramGenerator(q), a token generator that emits character q-grams. They are meant as encoders for document-vs-document tasks (classification, clustering, RandomIndexing into bit or multi-bit sketches), not for short-query search.

tc = TextConfig(tokenization=TokenizationConfig(nlist=[1],       # words; drop for q-grams only
        generators=[QgramGenerator(3), QgramGenerator(5)]))
voc = filter_tokens(t -> t.ndocs >= 3, Vocabulary(tc, corpus))   # prune the Zipfian tail
model = VectorModel(voc)
  • Q-grams are taken over the whole normalized text, boundary blanks included, so a window can span a word boundary. They do not span two fields of a multi-field document.
  • Runs of blanks count as one blank, without changing the normalized text that other generators read.
  • Tokens are tagged 'q'. The tag keeps the 3-gram que apart from the word que, and keeps word-level stopwords and lemmas from matching q-grams.
  • Q-grams use the ordinary Vocabulary, so VectorModel, EntropyWeighting, RandomIndexing, LSI and filter_tokens work unchanged. Because the filter_tokens predicate sees the token, words and q-grams can be pruned with separate thresholds by checking the tag.

Saving

  • A profile with q-gram generators is saved as format "1.2". A build that only reads "1.1" refuses it, instead of silently tokenizing without the q-grams.
  • Profiles without generators, the published ones included, still write "1.1" with the same bytes and the same profile_id.
  • Other custom generators still cannot be saved, and merge_profiles still rejects any config with generators.

Cost

A hash-keyed separate vocabulary was tried first. On 23,186 Markdown paragraphs with 16 threads, both versions built the same space: 314,775 distinct q-grams and the same number of nonzero entries.

build vocabulary vectorize memory
hash-keyed (dropped) 1.2 s 0.32 s 9 MB
QgramGenerator 7.5 s 0.86 s 19 MB

We kept the strings because a Dict is needed either way, and strings can be inspected. Most of the build-time gap comes from Vocabulary taking a lock per token (_locked_tokenize_and_push), which slows the word path equally. That is left for a separate change.

Test plan

  • Full test suite passes locally (Pkg.test()), including 33 new tests in test/testqgrams.jl: tokens, tagging, stopwords, pruning, entropy weighting, RandomIndexing, and a save/load round trip in both format versions.
  • Docs build (Documenter) not run.

🤖 Generated with Claude Code

QgramGenerator(q) emits character q-grams over the whole normalized text,
blanks included, so a window can span a word boundary; runs of blanks read as
one, and no q-gram spans two fields of a multi-field document. Tokens are
tagged 'q', which keeps the 3-gram "que" apart from the word "que" and keeps
word-level stopwords and lemmas off q-grams. Several lengths are several
generators, and they mix freely with word tokens.

They live on the ordinary Vocabulary, so VectorModel, EntropyWeighting,
RandomIndexing, LSI and filter_tokens work unchanged. A hash-keyed separate
vocabulary was tried first and built the identical space (314,775 q-grams on
23k Markdown paragraphs) about 6x faster, but a Dict is needed either way and
keeping the strings allows inspection; most of the gap is Vocabulary taking a
lock per token, which the word path pays too.

Profiles with q-gram generators save as format "1.2", so a "1.1"-only build
refuses them instead of silently tokenizing without the q-grams. Profiles
without generators, the published ones included, still write "1.1" with the
same bytes and profile_id. Other custom generators still refuse to save, and
merge_profiles still declines any generator list.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@sadit
sadit merged commit 71a9c16 into main Sep 24, 2026
3 of 5 checks passed
@sadit sadit mentioned this pull request Sep 25, 2026
sadit referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant