You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Do we need to suggest language-specific tokenisation (and provide details as to when to treat hyphen as a word character)?
Do we need to handle e.g. U+2010 as well as ASCII hyphen (U+002D which is also used as a minus sign)? Similarly to apostrophe, we ideally want to avoid having charset-specific variants of an algorithm (we used to and they got out of step in some cases) so probably the answer is to treat characters that can't be encoded in the specified character set as characters we won't see, at least in some cases.
Thinking about English (both what I know as a native speaker and https://en.wikipedia.org/wiki/Hyphen#Use_in_English) I think the best general rule there is to treat hyphen as a word break. That's not always what's wanted, but seems likely to cause fewer problems than treating it as a word character, or as a word joiner which is itself omitted from the word.
Possibly single-letter components could get special treatment (e.g. e-mail -> email, x-ray -> xray). I can't immediately think of a case where that goes wrong, but it's perhaps too much of a quirk for general tokenisation advice.
Split out of #187:
To do:
'h-' 'n-' 't-' //nAthair -> n-athair, but alone are problematic)нибудь | indef. suffix preceded by hyphen)Algorithms not annotated above currently have no special treatment of hyphens.