Class ContrastCodingTest

java.lang.Object
ubic.gemma.core.util.math.linearmodels.ContrastCodingTest

public class ContrastCodingTest extends Object
What a factor's columns look like under each ContrastCoding, and what its coefficients therefore mean.

The design matrix is the whole of the difference: both codings give a k-level factor k-1 columns, so the fit, the rank and everything downstream keep their shape, and only the numbers in those columns change. Most of these read the columns directly, because that is where the change is; the last two run a fit, because what the columns are FOR is what the coefficients then mean, and only a fit can show that.

  • Constructor Details

    • ContrastCodingTest

      public ContrastCodingTest()
  • Method Details

    • treatmentCodingLeavesTheDroppedLevelAtZero

      @Test public void treatmentCodingLeavesTheDroppedLevelAtZero()
      Treatment coding: the dropped level is all zeros, so each coefficient is a difference from it.
    • sumToZeroCodingPutsMinusOneInEveryColumnForTheDerivedLevel

      @Test public void sumToZeroCodingPutsMinusOneInEveryColumnForTheDerivedLevel()
      🛑 Sum-to-zero coding: the dropped level takes -1 in EVERY column of its factor. That is the whole mechanism — it is what makes the intercept the mean of the level means rather than the dropped level's mean, and each coefficient a deviation from that rather than a difference from a nominated control.

      The column count is unchanged, which is the other half of the point: nothing downstream of the design matrix has to know which coding was used in order to keep working.

    • codingIsChosenPerFactorNotForTheWholeDesign

      @Test public void codingIsChosenPerFactorNotForTheWholeDesign()
      The coding is per factor, not per matrix — a genotype factor with a real wild-type control keeps treatment coding while a tissue factor in the same design is deviation-coded. That combination is the reason the setting is not one flag on the analysis.
    • baselineAndCodingCommute

      @Test public void baselineAndCodingCommute()
      Setting the baseline after the coding has to give the same design as setting it before — both rebuild, and a caller has no reason to know the order matters. If it did, the level taking -1 would depend on the order two unrelated setters were called in, which is the kind of thing that is only ever found in a result.
    • sumToZeroCoefficientsAreDeviationsFromTheMeanOfTheLevelMeans

      @Test public void sumToZeroCoefficientsAreDeviationsFromTheMeanOfTheLevelMeans()
      🛑 The end of the contract, through an actual fit: under SUM_TO_ZERO the intercept is the mean of the LEVEL means and each coefficient is that level's deviation from it.

      The groups are deliberately unequal in size — 2, 4 and 1 sample — because on a balanced design the mean of the level means and the mean of the samples coincide, and a test built on one could not tell which reference the coefficients were against. Here they are 21.333 and 20.571.

      Level means 11, 23 and 30; their mean 64/3. Cortex is the derived level, so its deviation has no coefficient of its own and is minus the sum of the other two — which is the number the caller has to reconstruct, and the reason the analyzer synthesizes a contrast for it rather than leaving it absent.

    • treatmentCoefficientsAreDifferencesFromTheBaselineLevel

      @Test public void treatmentCoefficientsAreDifferencesFromTheBaselineLevel()
      The same data under TREATMENT coding, so the difference is visible rather than asserted. Here the intercept is cortex's own mean and every coefficient is a difference from cortex — which is the right answer when cortex is a control and the wrong question when it is just one tissue of three.
    • theDerivedLevelsColumnsAreFoundAndAreOnlyItsOwnFactors

      @Test public void theDerivedLevelsColumnsAreFoundAndAreOnlyItsOwnFactors()
      The derived level's columns are found, and they are exactly its own factor's.

      This is the guard against the silent half of DesignMatrix.getDerivedLevelColumns(): it declines to report a factor whose columns do not come out at one per non-derived level, so a naming change would turn the whole feature off rather than compute it wrongly. If that ever happens this test is what says so.

    • theDerivedLevelsContrastMatchesEstimatingItDirectly

      @Test public void theDerivedLevelsContrastMatchesEstimatingItDirectly()
      🛑 The synthesized contrast has to be the SAME contrast, standard error included.

      Which level is derived is an arbitrary choice, so fitting the same data with cortex derived and then with liver derived must give cortex the same deviation, standard error, t and p either way — estimated in one fit and reconstructed in the other. This is the check that the reconstruction is a real contrast and not a number that merely has the right sign.

      It is the standard error that this actually tests. The coefficient is minus the sum of the others by construction and would come out right under any arithmetic; the error requires the whole covariance block 1' XtXi 1. Summing the individual variances instead — the obvious wrong version — ignores the covariance between the estimates, which is not zero and has no fixed sign: on this design it is negative, so the wrong version comes out too LARGE (0.898 against 0.601). It would be too small on another design, which is why the assertion below pins "different" rather than a direction.

    • everyLevelGetsAContrastUnderSumToZero

      @Test public void everyLevelGetsAContrastUnderSumToZero()
      Every level of a sum-coded factor gets a row, including the one with no column.
    • treatmentCodingStillHasNoRowForTheBaseline

      @Test public void treatmentCodingStillHasNoRowForTheBaseline()
      Treatment coding is untouched: the baseline still has no row, because there is no contrast for it — it is what everything else is measured against, not a thing measured against something.
    • sumToZeroMatchesRContrSum

      @Test public void sumToZeroMatchesRContrSum()
      🛑 Benchmarked against R. The fits that produced every number below are in src/test/resources/data/stat-tests/contrast-coding.R, which regenerates them and carries the output as comments; these tests depend on them not changing.

      The level order there is not the obvious one and the script says why: R's contr.sum drops the LAST level, Gemma drops the FIRST — the one setBaseline names — so R's factor puts Gemma's baseline last to compare like with like.

      ⚠️ Gemma's "Std. Error" column holds the UNSCALED value (sqrt(diag(cov.unscaled))), not the standard error R prints. They differ by sigma, and the estimate, t and p are directly comparable while that one is not. Both forms are asserted here so the difference is recorded rather than rediscovered.

    • treatmentCodingMatchesRContrTreatment

      @Test public void treatmentCodingMatchesRContrTreatment()
      Treatment coding against R, so the sum-to-zero work is shown not to have disturbed the path everything else uses. Last block of contrast-coding.R; R's default, no contrasts() call needed.
    • anUnknownFactorIsRejected

      @Test public void anUnknownFactorIsRejected()
      Naming a factor that is not in the design is a caller error, not a silent no-op.