Repository navigation
among state machine enhancements #302
Description
Activity
I've been testing the new among code with wide characters - this requires an extra short in each state's entry, plus segments take up approximately twice as much space (more precisely, segments take up half as much space when not using wide characters, but rounded up if the segment length is odd). This is enough extra to push the
Step_2amonginserbian.sblover 32768 shorts, and we make use of the sign bit to distinguish a return value from an offset to another state in the table.(It might seem like this is a situation where the new approach is woefully inefficient, but it actually generates smaller tables for these cases: on a 64-bit platform where pointer members can be 4-byte aligned for UTF-8 stemmers the existing approach generates an among table that's 48840 bytes for the array plus 10765 (maybe plus padding) for the strings which is 59605 bytes total, while the new approach generates a 32810 byte table with no extra for strings; for wide characters using a 32-bit wchar_t it looks like the existing approach needs 48840 + 10765*4 = 91900 vs 78648 for the new approach).
The snowball compiler has long supported
-wwhen generating C/C++; libstemmer doesn't expose this (more in #267) but it's not helpful to regress this.The offsets to states are all currently from the start of the table, mostly as that's slightly simpler to execute. They could be from the current state, though that probably means we'd need to reorder states within the table (currently they're emitted using a depth-first approach). That would also make encoding
gopast non-groupinginto the table harder - for a grouping with multi-byte UTF-8 we'd then need to encode a transition to a state earlier in the table. The widechars tables could differ in this way, but that seems extra complexity better avoided.So I think one of the table size reducing ideas above is the best answer.
Merging common subtrees seems a good option. The leaf case looks fairly easy, and my quick hack at working out what it would save for serbian.sbl's among#2 gives 13234 words, which would reduce its among table to 32716 words, which is neatly just under the 32768 threshold.
Update: the non-leaf case is easy too - we can check the segment/N-way/2-way we've just encoded and if it exactly matches one we've already added we can shrink the table to where it started and return the offset to the exact match instead. Offsets in the encoded block matching means the sub-trees match too (a bit like git commit hashes being calculated over data including the hash of parent commits ensuring a match on commit hash means a match on ancestry). I have implemented locally for segments so far and everything works!
update2: Now merged to main.
I looked at the stats for sparse ranges. In UTF-8 and single-byte encodings, the codeunits are bytes so the maximum range size is 256 (and the first 32 are control characters in most of those encodings so unlikely to feature).
For wide characters we can get some much larger ranges though - here are all the ones >= 300 in size:
+++ NWAY window 1600:65276 144 of 63677 0.2%: 0x640 0x64b 0x64c 0x64d 0x64e 0x64f 0x650 0x651 0x652 0x660 0x661 0x662 0x663 0x664 0x665 0x666 0x667 0x668 0x669 0xfe80 0xfe81 0xfe82 0xfe83 0xfe84 0xfe85 0xfe86 0xfe87 0xfe88 0xfe89 0xfe8a 0xfe8b 0xfe8c 0xfe8d 0xfe8e 0xfe8f 0xfe90 0xfe91 0xfe92 0xfe93 0xfe94 0xfe95 0xfe96 0xfe97 0xfe98 0xfe99 0xfe9a 0xfe9b 0xfe9c 0xfe9d 0xfe9e 0xfe9f 0xfea0 0xfea1 0xfea2 0xfea3 0xfea4 0xfea5 0xfea6 0xfea7 0xfea8 0xfea9 0xfeaa 0xfeab 0xfeac 0xfead 0xfeae 0xfeaf 0xfeb0 0xfeb1 0xfeb2 0xfeb3 0xfeb4 0xfeb5 0xfeb6 0xfeb7 0xfeb8 0xfeb9 0xfeba 0xfebb 0xfebc 0xfebd 0xfebe 0xfebf 0xfec0 0xfec1 0xfec2 0xfec3 0xfec4 0xfec5 0xfec6 0xfec7 0xfec8 0xfec9 0xfeca 0xfecb 0xfecc 0xfecd 0xfece 0xfecf 0xfed0 0xfed1 0xfed2 0xfed3 0xfed4 0xfed5 0xfed6 0xfed7 0xfed8 0xfed9 0xfeda 0xfedb 0xfedc 0xfedd 0xfede 0xfedf 0xfee0 0xfee1 0xfee2 0xfee3 0xfee4 0xfee5 0xfee6 0xfee7 0xfee8 0xfee9 0xfeea 0xfeeb 0xfeec 0xfeed 0xfeee 0xfeef 0xfef0 0xfef1 0xfef2 0xfef3 0xfef4 0xfef5 0xfef6 0xfef7 0xfef8 0xfef9 0xfefa 0xfefb 0xfefc +++ NWAY window 32:8205 10 of 8174 0.1%: 0x20 0x623 0x624 0x625 0x626 0x629 0x643 0x64a 0x6c1 0x200d +++ NWAY window 108:539 6 of 432 1.4%: 0x6c 0x72 0x74 0x76 0x103 0x21b +++ NWAY window 99:539 10 of 441 2.3%: 0x63 0x6c 0x6e 0x72 0x73 0x74 0x76 0x103 0x219 0x21b +++ NWAY window 97:537 6 of 441 1.4%: 0x61 0x69 0x6e 0x73 0x75 0x219 +++ NWAY window 97:539 8 of 443 1.8%: 0x61 0x65 0x74 0x75 0x7a 0xe2 0x219 0x21bAll of these would be reduced a lot by allowing two ranges with a gap.
I've been working on a new approach to implementing Snowball's
substring...among.This essentially encodes a state machine where each transition is either an O(1) multi-way dispatch on the next byte/character or a check that a particular string of bytes/characters follows (in UTF-8 it works in bytes; for fixed-width encodings it works in characters). This makes among O(1) in the number of strings rather than O(log(n)), and testing shows it's actually faster in practice for real-world among use. The size of the data tables is typically smaller than the old implementation
on 64-bit machines, and roughly comparable on 32-bit machines. The generated C code is much more shared library friendly (the existing approach results in a lot of dynamic load time relocations).This is now working well, and I just need to clean it up a bit and then merge it.
Once that's done, there are some further enhancements that could be made which I'm opening this ticket to keep track of. In no particular order:
among ( 'ion' 'ions' 'ian' 'ians' ). This reduces the size of the data and so improves cache utilisation.Might need 2 passes, or some sort of interim data structure.min_length_matchin the code (which is calculated but currently not used).After a segment, the next transition can't be another segment - can we take advantage of that? Could encode segment+N-way with segment sometimes empty. Might work against common-subtree sometimes. Also a segment longer than 255 bytes/characters would probably need encoding as multiple consecutive segments...(Consecutive segments can occur if we can exactly match part way through - e.g. backwardsamong ( 'soyad' 'ad' )has segmentadthen segmentsoy.)find_among()/find_among_b()- that means we could inline their code like we do for explicit calls. A C/C++ compiler should be able to inline them for us now (whereas it couldn't for the old among implementation), but this inlining would help languages which don't have an optimiser which inlines code for us. Even for C/C++, it could enable turning some global snowball integer or boolean variables into locals.gopast non-GROUPINGcould be encoded as an among table where the state machine can loop (it consumes a byte/character on each step so any looping is finite). Alsogotoand non-inverted cases, but that would require us to support the default (a) being a state transition and (b) advancing the cursor so it a bit more complex.amongs could be encoded together into a single table (including an among gating function which is an action-less among); similarly for grouping checks in such positions.