Class QueryTokens

java.lang.Object
ubic.gemma.rest.ranking.QueryTokens

public final class QueryTokens extends Object
The one query tokeniser the annotation-search stack uses.

Moved here from AnnotationsWebService, which had the careful implementation, while CompositeRankingStrategy and TokenCoverageRankingStrategy each carried their own query.toLowerCase().split("\\s+"). That is worse than it sounds, because coverage is scored by substring containment: with no stop-word strip, the in "cell line of the liver" scores against theca cell and of against profile, handing free coverage to labels that share nothing but two letters. On a real gold pair, "epithelium of esophagus", the of token did exactly that.

  • Method Details

    • contentTokens

      public static List<String> contentTokens(@Nullable String query)
      Tokenise an arbitrary user query into "content" tokens: lowercase, split on runs of non-alphanumeric characters, drop tokens shorter than MIN_CONTENT_TOKEN_LENGTH, drop stop-words.

      Returned in encounter order, deduplicated; empty list when the input is null / blank / all-stop-words. Callers should treat an empty list as "no token-coverage constraint applies — fall back to Lucene's order".