Class ExpressionExperimentDaoTest

java.lang.Object
ubic.gemma.core.util.test.BaseTest5
ubic.gemma.core.util.test.BaseDatabaseTest5
ubic.gemma.persistence.service.expression.experiment.ExpressionExperimentDaoTest

@ContextConfiguration @TestExecutionListeners(value=org.springframework.security.test.context.support.WithSecurityContextTestExecutionListener.class, mergeMode=MERGE_WITH_DEFAULTS) public class ExpressionExperimentDaoTest extends BaseDatabaseTest5
  • Constructor Details

    • ExpressionExperimentDaoTest

      public ExpressionExperimentDaoTest()
  • Method Details

    • removeFixtures

      @AfterEach public void removeFixtures()
    • loadTroubledIds

      @Test public void loadTroubledIds()
    • testGetFilterableProperties

      @Test public void testGetFilterableProperties()
    • testGetFilterWithStatementObject

      @Test public void testGetFilterWithStatementObject()
    • testFilterWithStatementSecondObject

      @Test @WithMockUser public void testFilterWithStatementSecondObject()
    • testDatasetsCanBeFilteredByVisibility

      @Test @WithMockUser public void testDatasetsCanBeFilteredByVisibility()
      🛑 `isPublic` is not a mapped attribute — it is read off the ACL entries when the value object is built — so the metamodel walk that enumerates filterable properties could not see it and `filter=isPublic = false` answered 400 "the property of isPublic is unknown". It is now registered by hand and resolved to the ACL predicate.

      This asserts the query RUNS: the expression is a case-over-exists carrying the ACL class-id parameter, and the ways it can be wrong are all failures of the HQL rather than wrong answers — an unbound parameter, or a property Hibernate cannot resolve. A filter that never executed would be no better than no filter at all.

    • testThawTransientEntity

      @Test public void testThawTransientEntity()
    • testThaw

      @Test public void testThaw()
    • testThawLite

      @Test public void testThawLite()
    • testThawLiter

      @Test public void testThawLiter()
    • testLoadWithRefreshCacheMode

      @Test public void testLoadWithRefreshCacheMode()
    • testLoadReference

      @Test public void testLoadReference()
    • testRemovingAQuantitationTypeFromItsExperimentDeletesTheRow

      @Test public void testRemovingAQuantitationTypeFromItsExperimentDeletesTheRow()
      Taking a quantitation type out of its experiment deletes the row.

      Without orphanRemoval the join column is set to NULL and the row stays behind, owned by no experiment and unreachable from every finder, all of which take an experiment.

    • testLoadMultipleReferences

      @Test public void testLoadMultipleReferences()
    • testGetTechnologyTypeUsageFrequency

      @Test @WithMockUser public void testGetTechnologyTypeUsageFrequency()
    • testGetArrayDesignUsageFrequency

      @Test @WithMockUser public void testGetArrayDesignUsageFrequency()
    • testGetOriginalPlatformUsageFrequency

      @Test @WithMockUser public void testGetOriginalPlatformUsageFrequency()
    • testGetCategoriesWithUsageFrequency

      @Test @WithMockUser(authorities="GROUP_ADMIN") public void testGetCategoriesWithUsageFrequency()
    • testGetCategoriesUsageFrequencyAsAnonymous

      @Test @WithMockUser public void testGetCategoriesUsageFrequencyAsAnonymous()
    • testGetCategoriesUsageFrequencyWithIds

      @Test public void testGetCategoriesUsageFrequencyWithIds()
      No ACL filtering is done when explicit IDs are provided, so this should work without WithMockUser.
    • testGetAnnotationUsageFrequency

      @Test @WithMockUser(authorities="GROUP_ADMIN") public void testGetAnnotationUsageFrequency()
    • testGetAnnotationUsageFrequencyAsAnonymous

      @Test @WithMockUser public void testGetAnnotationUsageFrequencyAsAnonymous()
    • testGetAnnotationUsageFrequencyWithLargeBatch

      @Test @WithMockUser(authorities="GROUP_ADMIN") public void testGetAnnotationUsageFrequencyWithLargeBatch()
    • testGetAnnotationUsageFrequencyRetainMentionedTerm

      @Test @WithMockUser(authorities="GROUP_ADMIN") public void testGetAnnotationUsageFrequencyRetainMentionedTerm()
    • testGetAnnotationUsageFrequencyExcludingFreeTextTerms

      @Test @WithMockUser(authorities="GROUP_ADMIN") public void testGetAnnotationUsageFrequencyExcludingFreeTextTerms()
    • testGetAnnotationUsageFrequencyExcludingFreeTextCategories

      @Test @WithMockUser(authorities="GROUP_ADMIN") public void testGetAnnotationUsageFrequencyExcludingFreeTextCategories()
    • testGetAnnotationUsageFrequencyExcludingUncategorizedTerms

      @Test @WithMockUser(authorities="GROUP_ADMIN") public void testGetAnnotationUsageFrequencyExcludingUncategorizedTerms()
    • testGetAnnotationUsageFrequencyWithUncategorizedCategory

      @Test @WithMockUser(authorities="GROUP_ADMIN") public void testGetAnnotationUsageFrequencyWithUncategorizedCategory()
    • testGetAnnotationUsageFrequencyWithIds

      @Test public void testGetAnnotationUsageFrequencyWithIds()
      No ACL filtering is done when explicit IDs are provided, so this should work without WithMockUser.
    • testGetPerTaxonCount

      @Test @WithMockUser("bob") public void testGetPerTaxonCount()
    • testFilterAndCountByArrayDesign

      @Test @WithMockUser public void testFilterAndCountByArrayDesign()
    • testSubquery

      @Test public void testSubquery()
    • testSubqueryWithMultipleJointures

      @Test public void testSubqueryWithMultipleJointures()
    • testGetSubSetsByExpressionExperimentsEmpty

      @Test public void testGetSubSetsByExpressionExperimentsEmpty()
    • testGetSubSetsByExpressionExperimentsBatched

      @Test public void testGetSubSetsByExpressionExperimentsBatched()
    • testRemoveExperimentWithSharedBioMaterial

      @Test public void testRemoveExperimentWithSharedBioMaterial()
    • removeWithBioAssayDimension

      @Test public void removeWithBioAssayDimension()
    • testLoadValueObjectNamesThePlatformsAndTheSwitch

      @Test @WithMockUser public void testLoadValueObjectNamesThePlatformsAndTheSwitch()
      The dataset VO has to name the platform, not merely count it: a curator reading a dataset in the UI sees a blank platform line otherwise, and originalPlatform is the field that says a dataset was switched, which is a thing they check for. Asked for by uib with Paul's "as long as it is fast to fetch" — it is the same join the array-design COUNT was already making, so nothing extra is fetched.
    • testLoadValueObjectReportsNoSwitchWhenThereWasNone

      @Test @WithMockUser public void testLoadValueObjectReportsNoSwitchWhenThereWasNone()
      The other half: a dataset nobody switched reports no original platform at all. Without this a report that named every used platform as an original one would pass the test above.
    • testTechnologyTypeIsNullWhenThePlatformsDisagree

      @Test @WithMockUser public void testTechnologyTypeIsNullWhenThePlatformsDisagree()
      🛑 A dataset run on two kinds of platform IS both, so it has no single technology and the field says so by being null. Answering with either platform's type — which is what the details VO does, taking whichever platform the iterator reaches first — labels half the dataset wrong, and a client cannot tell that from a confident answer.
    • testHasSourceMetadata

      @Test @WithMockUser public void testHasSourceMetadata()
      The backfill asks this once per experiment across the corpus to decide whether to skip it, so it must not read the document to answer — SOURCE_METADATA is a LONGTEXT holding the whole GEO record.
    • testDateCreatedComesFromTheCreationAuditEvent

      @Test @WithMockUser public void testDateCreatedComesFromTheCreationAuditEvent()
    • testLoadValueObjectWithSingleCellData

      @Test @WithMockUser public void testLoadValueObjectWithSingleCellData()
    • testLoadValueObjectCarriesTheOtherPartsOfASplitStudy

      @Test @WithMockUser public void testLoadValueObjectCarriesTheOtherPartsOfASplitStudy()
      A split part names its siblings on the dataset itself.

      Gemma splits an experiment by a factor and titles each part "Split part N of: …", which tells a reader siblings exist and gives no way to reach one — 52 of 100 sampled single-cell datasets are split parts (uib, 2026-09-03), and the field lived on a VO only /experiment-sets/{id}/datasets serves.

    • testLoadValueObjectSummarizesTheLibraryFieldsAcrossSamples

      @Test @WithMockUser public void testLoadValueObjectSummarizesTheLibraryFieldsAcrossSamples()
      The dataset says what kind of libraries its samples were made from, so a client does not have to fetch the sample list to read one constant.

      technologyType cannot answer it: Gemma maps sequencing onto generic gene-list platforms, so GSE270825 reads GENELIST with 24 SSRNA_SEQ samples (uib, 2026-09-16). Reading it off /datasets/{id}/samples instead cost 652 KiB gzipped and 2.3 s for GSE2109's 2,158 assays, for one line of text on a page.

    • testLoadValueObjectCountsSamplesWithNoValueUnderANullEntry

      @Test @WithMockUser public void testLoadValueObjectCountsSamplesWithNoValueUnderANullEntry()
      🛑 A sample carrying no value is counted under a null-valued entry, not dropped.

      Dropping it would break the sum against numberOfBioAssays and would render a microarray dataset's librarySelection — null on every assay, because the technology has no selection step — as an empty list, which reads as "this dataset has no samples" rather than "no value".

    • testLoadValueObjectLibraryFieldsAreEmptyForADatasetWithNoSamples

      @Test @WithMockUser public void testLoadValueObjectLibraryFieldsAreEmptyForADatasetWithNoSamples()
      A dataset with no samples gets empty lists, not null — one shape for every dataset.
    • testLoadValueObjectOtherPartsIsEmptyForAnUnsplitDataset

      @Test @WithMockUser public void testLoadValueObjectOtherPartsIsEmptyForAnUnsplitDataset()
      An unsplit dataset gets an empty list, not null — one shape for every dataset.
    • testLoadValueObjectWithoutSingleCellDataIsNotSingleCell

      @Test @WithMockUser public void testLoadValueObjectWithoutSingleCellDataIsNotSingleCell()
      And an ordinary dataset says so, rather than leaving the client to infer it.
    • removeWithUninitializedSingleCellDataVectors

      @Test @WithMockUser public void removeWithUninitializedSingleCellDataVectors()
      Removing an experiment whose single-cell vectors were never loaded must go through the projection-query branch of removeAllSingleCellDataVectors rather than walking the lazy collection. Walking it selects every vector's DATA + DATA_INDICES blob: on GSE277430 (25,050 vectors over 333,570 cells) that exhausted a 30 GB heap, and the resulting OutOfMemoryError desynced the JDBC connection so the failure surfaced as an ArrayIndexOutOfBoundsException thrown during rollback, with the real cause discarded as "Application exception overridden by rollback exception".

      Every other test in this area builds its fixture in-session, so the collection is already initialized and only the other branch runs. This one reaches the branch that runs in production and therefore also validates that its HQL parses.

    • testGetAllAnnotations

      @Test public void testGetAllAnnotations()
    • testGetAnnotationsByLevel

      @Test public void testGetAnnotationsByLevel()
    • testGetRawDataVectors

      @Test public void testGetRawDataVectors()
    • testAddRawDataVectors

      @Test public void testAddRawDataVectors()
    • testAddRawDataVectorsWithNumberOfCells

      @Test public void testAddRawDataVectorsWithNumberOfCells()
    • testRemoveAllRawDataVectors

      @Test public void testRemoveAllRawDataVectors()
    • testRemoveRawDataVectors

      @Test public void testRemoveRawDataVectors()
    • testRemoveRawDataVectorsWithDimension

      @Test public void testRemoveRawDataVectorsWithDimension()
    • testRemoveRawDataVectorsWhenQtIsUnknown

      @Test public void testRemoveRawDataVectorsWhenQtIsUnknown()
    • testRemoveRawDataVectorsWithNumberOfCells

      @Test public void testRemoveRawDataVectorsWithNumberOfCells()
    • testReplaceRawDataVectors

      @Test public void testReplaceRawDataVectors()
    • testReplaceRawDataVectorsWithNewDimension

      @Test public void testReplaceRawDataVectorsWithNewDimension()
    • testCreateProcessedDataVectors

      @Test public void testCreateProcessedDataVectors()
    • testCreateProcessedDataVectorsWhenCollectionIsNotInitialized

      @Test public void testCreateProcessedDataVectorsWhenCollectionIsNotInitialized()
      An experiment whose processed vectors collection was never initialized is checked for existing vectors by query, and its collection is left uninitialized after the vectors are persisted, loading them from the database when accessed.
    • testCreateProcessedDataVectorsInSeveralBatchesWithNumberOfCells

      @Test public void testCreateProcessedDataVectorsInSeveralBatchesWithNumberOfCells()
      More vectors than fit in a single batch, each with a number of cells: every vector and its number of cells must be persisted.
    • testCreateProcessedDataVectorsWithNonPersistentQt

      @Test public void testCreateProcessedDataVectorsWithNonPersistentQt()
    • testRemoveProcessedDataVectors

      @Test public void testRemoveProcessedDataVectors()
    • testReplaceProcessedDataVectors

      @Test public void testReplaceProcessedDataVectors()
    • testReplaceProcessedDataVectorsReusingTheSameQT

      @Test public void testReplaceProcessedDataVectorsReusingTheSameQT()
    • testReplaceProcessedDataVectorsWithDetachedExperiment

      @Test public void testReplaceProcessedDataVectorsWithDetachedExperiment()
    • testGetGenesUsedByProcessedVectors

      @Test public void testGetGenesUsedByProcessedVectors()
    • testGetArrayDesignUsed

      @Test public void testGetArrayDesignUsed()
    • testFindIdsByBioMaterial_resolvesThroughASubset

      @Test @WithMockUser(authorities="GROUP_ADMIN") public void testFindIdsByBioMaterial_resolvesThroughASubset()
      A sample whose only route to an experiment is subset -> sourceExperiment must still resolve.

      The subset fallback in findIdsByBioMaterial was guarded on results == null, and Query.list() returns an empty list rather than null, so it never ran once. Every aggregated single-cell sample resolved to nothing: 665,120 of the 669,233 samples an ACL repair had to parent on 2026-08-30, each one logging "Could not find an ExpressionExperiment associated to BioMaterial".

    • testUpdateMeanVarianceRelation_deletesTheOneItReplaces

      @Test @WithMockUser(authorities="GROUP_ADMIN") public void testUpdateMeanVarianceRelation_deletesTheOneItReplaces()
      Replacing an experiment's mean-variance relation must delete the one it replaces.

      ExpressionExperiment.meanVarianceRelation is a @ManyToOne, and JPA has no orphanRemoval for one, so moving the reference used to leave the previous row behind with nothing pointing at it. Nothing sweeps those: on production 2026-08-30, 33,535 of 57,311 rows were unreferenced, roughly 15 GB of the table's 25.4 GB, since each row carries four mediumblobs. Abandoned ids ran right up to the maximum, so it was still accumulating.

    • testGetSubSetsToleratesATransientBioAssayDimension

      @Test public void testGetSubSetsToleratesATransientBioAssayDimension()
      🛑 A TRANSIENT BioAssayDimension is a normal input here, and it must not throw.

      DiffExAnalyzerUtils.dropSamplesNotAnalyzed re-slices the data matrix whenever a sample is DE_Exclude or an outlier, and the replacement dimension comes from createBADMap -> BioAssayDimension.Factory.newInstance, which is never persisted. The old query bound the dimension entity (... and bad = :bad ...), so an unsaved one raised TransientObjectException: object references an unsaved transient instance ... BioAssayDimension on the subset REUSE LOOKUP — before any write, which is why the transaction rolled back intact.

      frinkbro hit it on GSE62625 (eid 9439) on 2026-09-10 and held 18 further subset re-runs. It bit only some experiments because dropSamplesNotAnalyzed returns the original matrix, and so the persisted dimension, when nothing is dropped.

    • testGetSubSetsStillReusesWhenASampleWasDropped

      @Test public void testGetSubSetsStillReusesWhenASampleWasDropped()
      Reuse still works when a sample has been DROPPED — the case that decides between the two possible fixes.

      Skipping the lookup for a transient dimension would also stop the crash, and would silently build a duplicate subset beside the one the experiment already has, on every experiment carrying a DE_Exclude marker — 345 of them, 326 under collection of material (frinkbro, 2026-09-10). Matching on the dimension's ASSAYS keeps reuse working: a subset whose assays all survive the drop is still found, and one that contained the dropped assay is correctly NOT found, because it can no longer be fully covered.