Class QuantileMissingValueMaskingTest

java.lang.Object
ubic.gemma.core.util.math.QuantileMissingValueMaskingTest

public class QuantileMissingValueMaskingTest extends Object
Which cells come back missing from quantile normalization.

A missing cell is imputed with its row mean so it can take part in the ranking, then re-masked at the end. The record of which cells to re-mask is a BitSet over row * columns + column rather than a same-shaped matrix, and an index that is off — transposed, or using the wrong dimension as the divisor — puts the NaNs back in the wrong places while every other value stays correct. These tests are what would catch that.

  • Constructor Details

    • QuantileMissingValueMaskingTest

      public QuantileMissingValueMaskingTest()
  • Method Details

    • missingCellsAreRemaskedInTheirOwnPositions

      @Test public void missingCellsAreRemaskedInTheirOwnPositions()
      🛑 Non-square and asymmetric on purpose.

      On a square matrix a transposed index is invisible, and with a whole column missing — the case QuantileReferenceColumnsTest exercises — any row-major/column-major mix-up still lands inside the same column. Neither shape can fail. Three rows by five columns, with the two missing cells in different rows AND different columns, is the smallest thing that can.

    • completeDataComesBackComplete

      @Test public void completeDataComesBackComplete()
      Data with nothing missing comes back with nothing missing. The empty case is worth pinning separately because it is what the whole corpus looks like — the raw vectors measured on GSE260875 have no missing values in any of their three quantitation types — so a masking bug that only fires on a set bit would never show up in production and would still be wrong.
    • anAllMissingRowIsDroppedAndTheRestKeepTheirNames

      @Test public void anAllMissingRowIsDroppedAndTheRestKeepTheirNames()
      A row that is missing everywhere is dropped rather than returned as a row of NaN — RowMissingFilter with minPresentCount=1 removes it before anything else runs, so the result is shorter than the input and the surviving rows keep their names. The masking has to be indexed on the FILTERED shape; using the input's column or row count would run off the end or mask the wrong cells.