Class QuantileNormalizer<R,C>

java.lang.Object
ubic.gemma.core.analysis.preprocess.normalize.QuantileNormalizer<R,C>

public class QuantileNormalizer<R,C> extends Object
Perform quantile normalization on a matrix, as described in:

Bolstad, B (2001) _Probe Level Quantile Normalization of High Density Oligonucleotide Array Data_. Unpublished manuscript PDF

Bolstad, B. M., Irizarry R. A., Astrand, M, and Speed, T. P. (2003) _A Comparison of Normalization Methods for High Density Oligonucleotide Array Data Based on Bias and Variance._ Bioinformatics 19(2) ,pp 185-193. web page. However, note that this deals with missing values differently than the Bioconductor implementation.
Author:
pavlidis
See Also:
  • Constructor Details

    • QuantileNormalizer

      public QuantileNormalizer()
  • Method Details

    • normalize

      public DoubleMatrix<R,C> normalize(DoubleMatrix<R,C> dataMatrix)
    • normalize

      public DoubleMatrix<R,C> normalize(DoubleMatrix<R,C> dataMatrix, @Nullable boolean[] includeInReference)
      Normalize, letting only some columns define the reference distribution.
      Parameters:
      includeInReference - one flag per column; null means every column contributes
      See Also:
    • normalizeInPlace

      public static void normalizeInPlace(double[][] data, @Nullable boolean[] includeInReference)
      Normalize rows of data in place, producing exactly the values of normalize(DoubleMatrix, boolean[]).

      The matrix path holds the input matrix, a filtered copy of it and a sorted copy of that, three full-size rows x columns copies on top of the caller's own data. This one works on the caller's row arrays and needs two double[rows] buffers, a double[rows] reference distribution and one bit per cell of the rows that have missing values.

      Every step is the same computation in the same order, which is what makes the results identical rather than merely close: rows with no value at all are left alone (the matrix path drops them with RowMissingFilter), a missing cell is imputed with its row mean for the ranking, each column is sorted with AbstractList.sort(), the reference at each rank is the mean over the contributing columns summed in column order, ranks come from Rank.rankTransform(DoubleArrayList) with the same tie handling, and the missing cells are masked again at the end.

      Parameters:
      data - one array per row, all of the same length; rows that have at least one value are overwritten with their normalized values, rows with none are left as they are
      includeInReference - one flag per column; null means every column contributes
      See Also: