Class QueryTokens
java.lang.Object
ubic.gemma.rest.ranking.QueryTokens
The one query tokeniser the annotation-search stack uses.
Moved here from AnnotationsWebService, which had the careful implementation, while
CompositeRankingStrategy and TokenCoverageRankingStrategy each carried their own
query.toLowerCase().split("\\s+"). That is worse than it sounds, because coverage is
scored by substring containment: with no stop-word strip, the in
"cell line of the liver" scores against theca cell and of against
profile, handing free coverage to labels that share nothing but two letters. On a real
gold pair, "epithelium of esophagus", the of token did exactly that.
-
Method Summary
Modifier and TypeMethodDescriptioncontentTokens(String query) Tokenise an arbitrary user query into "content" tokens: lowercase, split on runs of non-alphanumeric characters, drop tokens shorter thanMIN_CONTENT_TOKEN_LENGTH, drop stop-words.
-
Method Details
-
contentTokens
Tokenise an arbitrary user query into "content" tokens: lowercase, split on runs of non-alphanumeric characters, drop tokens shorter thanMIN_CONTENT_TOKEN_LENGTH, drop stop-words.Returned in encounter order, deduplicated; empty list when the input is null / blank / all-stop-words. Callers should treat an empty list as "no token-coverage constraint applies — fall back to Lucene's order".
-