Dr. SheikhBeig's workflow is now shown from A to Z.
The thesis is following the M1/M2 pathway Dr. SheikhBeig showed: monocyte/M0 → M1 and monocyte/M0 → M2 are treated as two branches, then compared. Right now we are still at the data-foundation stage: databases are found, dataset roles are proposed, and the evidence package is waiting for Dr. SheikhBeig's review.
What pathway?
Dr. SheikhBeig's M1/M2 workflow
Two branches: M0/monocyte → M1 and M0/monocyte → M2, then compare.What is done?
Database step is done
Confirmed internally and shown here. Relevant public resources were found.Where are we now?
Data checks are waiting for review
Step 2A–2D checks are done internally and waiting for review by Dr. SheikhBeig.What is not done?
No analysis yet
No preprocessing, no network, no regulator discovery, no final candidates.The whole workflow, read from top to bottom
One shared starting point splits into two parallel branches. Each branch runs down the same stages, and the two branches only come back together at the final candidate list. Status colors show where we are; the path itself does not change with progress.
Monocyte / M0
Shared starting point of both branches.
Monocyte/M0 → M1
Classical activation, usually IFNγ + LPS.
Monocyte/M0 → M2
Alternative activation, usually IL-4/IL-13 or IL-10.
Collect public records (GEO/SRA and support databases) holding M0 → M1 samples.
Collect public records (GEO/SRA and support databases) holding M0 → M2 samples.
Check donors, sample labels, matrices, and papers behind the M1 arm.
Check donors, sample labels, matrices, and papers behind the M2 arm.
Reprocess raw reads, build count matrices, confirm M1 markers.
Reprocess raw reads, build count matrices, confirm M2 markers.
Join mRNA with miRNA and methylation layers, then build the M1 network.
Join mRNA with miRNA and methylation layers, then build the M2 network.
Rank M1 hubs and transcription-factor regulators.
Rank M2 hubs and transcription-factor regulators.
GO/KEGG/Reactome meaning of the M1 modules.
GO/KEGG/Reactome meaning of the M2 modules.
Final candidate outputs
Shared regulators, M1-specific candidates, and M2-specific candidates, graded by evidence. This has not been produced yet.
Data and integration checks requested by Dr. SheikhBeig
The points Dr. SheikhBeig raised are grouped into a short set of confirmation checks below, each with what was checked, what the result was, and the decision taken from it.
GSE117040 (RNA-seq) with GSE117124 (methylation).| Confirmation check | Result | Decision |
|---|---|---|
| Is the data actually available? | Yes. GEO/NCBI, SRA, and PubMed were searched, with ENCODE, BLUEPRINT, and the Human Cell Atlas as reference resources. Usable public M1/M2 macrophage datasets were found, and each candidate has a defined role. | The data-finding step is complete; work moves to how the datasets are used. |
| Which omics layers are usable? | Three core layers: RNA-seq, DNA methylation (MBD-seq), and miRNA-seq. lncRNA may be extracted from RNA-seq where the annotation supports it. | Report three core omics layers; describe lncRNA as RNA-seq-derived, not as a separate measured layer. |
| Are the datasets biologically and sample-wise compatible? | Primary human macrophage, PBMC-derived macrophage, and THP-1 records were separated, and M0/M1/M2 structure was confirmed. GSE162698 is the strongest primary-human backbone (M0 plus M1 and M2 endpoints, donor structure, raw SRA data); GSE117040 is independent M1/M2 RNA-seq; GSE117124 shares its study framework and donor labels A–D. THP-1 lines, siRNA/transfected samples, and the unmatched miRNA cohort are not comparable to primary untreated human material. |
Use GSE162698 as the primary discovery backbone and GSE117040 as independent validation, kept separate rather than pooled. GSE117124 with GSE117040 is the compatible methylation pair. Non-comparable records stay support-only. |
| Can everything be integrated directly, and can batch effects be handled? | No to full direct integration. Cell source, donor structure, perturbation, platform, and context differ, so the datasets are not one fully matched multi-omics cohort. Technical differences such as laboratory and platform can be modelled, but biological differences between non-comparable sources should not be corrected away. | Use an evidence-layered integration design instead of claiming a matched cohort, and correct technical batch effects only where the underlying biology is comparable. |
| Where does mixOmics apply? | Mainly to GSE117040 with GSE117124, where sample and design compatibility holds. It is weak for unrelated or unmatched datasets. |
Use mixOmics only for compatible blocks; use evidence-layered integration elsewhere. |
| What is the current boundary and the next gate? | The stack covers general human macrophage M1/M2 polarization, not one shared disease-specific cohort. No network or regulator analysis has started. | Make no regulator claims; this stage covers data and integration feasibility only. Disease-specific context stays a later extension unless compatible datasets are added. |
| Role | Dataset / resource | Use |
|---|---|---|
| Primary discovery backbone | GSE162698 | Main M0 → M1 and M0 → M2 RNA-seq analysis. |
| Independent RNA-seq validation | GSE117040 | Validate M1-vs-M2 signatures and candidate regulators. |
| Methylation support and direct integration candidate | GSE117124 with GSE117040 | Best RNA + methylation integration candidate; possible mixOmics use. |
| miRNA support | GSE51307 | Post-transcriptional support layer, not a directly matched integration layer unless matching is shown. |
| Perturbation and context support | GSE329556 | Support and context only, because of the siRNA/transfection design and the absence of raw data. |
| Prior-art benchmark | GSE46903 / GSE47189 | Compare against the known macrophage activation literature. |
| Cell-line support only | GSE273627 / GSE273628 | THP-1 support only; not used for primary-human claims. |
| Reference and background | ENCODE / BLUEPRINT / HCA / CELLxGENE | Annotation and background only, unless a compatible matched dataset is selected. |
Public databases were found and assigned to the workflow
This part is already confirmed internally. The point was to answer: “Which public databases can support Dr. SheikhBeig's pipeline?”
GEO / NCBI
Main source for human macrophage RNA-seq, miRNA-seq, methylation, ATAC/ChIP, and supplementary matrices.
Main dataset sourceSRA / ENA
Raw sequencing reads behind GEO records. Needed when formal analysis must reprocess data rather than rely on TPM or normalized matrices.
Raw data sourceArrayExpress / BioStudies / OmicsDI
Cross-check source for older expression and multi-omics records not always visible through GEO-only search.
Cross-checkCELLxGENE / Human Cell Atlas
Single-cell/tissue atlas reference layer. Useful for projection/background, but not the controlled M1/M2 backbone.
Reference layerENCODE
Reference regulatory/chromatin data for monocytes and macrophages. Useful for annotation, not main M1/M2 contrast.
Annotation supportBLUEPRINT / EGA
Immune-cell epigenomic support layer; some resources are controlled access.
Epigenome supportPubMed / publisher records
Used to confirm each accession, paper, method, data availability, and prior art.
Paper confirmationDataset checks are done internally; the proposed stack is ready for review
This step answers whether the found datasets are actually usable: human or cell line, donor structure, M0/M1/M2 labels, matrix files, raw data, and publication support.
What was checked?
Step 2A–2D checked GEO/GSM metadata, matrix availability, donor and column mapping, and publication-level confirmation.
Inputs checked
- 647 GSM sample records
- 13 GEO series records
- 17 supplementary entries
Mapping checked
- GSM sample titles
- Matrix columns
- Donor labels
- M0/M1/M2 labels
Current decision
- Done internally
- Not approved by Dr. SheikhBeig yet
- No analysis starts before review
Proposed dataset stack
These are the datasets picked into roles. They are proposed roles, not final locked analysis roles.
GSE162698
Primary discovery proposal. Primary human monocyte-derived macrophages; 3 donors; M0, M1, M2-IL4, M2-IL10, TAM-like. Processed TPM exists; raw SRA exists.
Discovery candidateGSE117040
Independent RNA-seq validation. Donors A–D, paired M1/M2 RNA-seq. No M0 baseline, so it validates M1-vs-M2 rather than the full branch.
ValidationGSE117124
Methylation support. MBD-seq for the same donor-label structure as GSE117040; useful as an epigenetic support layer.
MethylationGSE329556
Perturbation-context support. Primary human MDMs, siNC vs siXYLT2 across M0/M1/M2. Not pooled with untreated datasets; raw data unavailable.
Support onlyGSE46903 / GSE47189
Prior-art benchmark. Xue et al. macrophage activation-spectrum resource. Important because it already contains network/regulator analysis.
Prior artGSE51307
miRNA support. miRNA-seq arm with M0/M1/M2/TPP samples. Support layer only; not donor-matched to the discovery backbone.
miRNAGSE273627 / GSE273628
THP-1 support only. These are cell-line records, not primary human donor datasets; they cannot carry discovery or validation claims.
Cell line onlyPreprocessing starts only after Dr. SheikhBeig approves the data stack
This is the next practical step, but it has not started for the M1/M2 stack.
What will happen here?
Download/reprocess raw SRA where available, create count matrices, perform QC, confirm sample labels, and verify known M1/M2 marker behavior before network work.
Rule
- Processed matrices are for preview/QC only
- Formal DE/network work should use raw reprocessing where possible
Exception
- GSE329556 has no raw data
- Use only as support/context
Status
- Not started
- Waiting for Dr. SheikhBeig review of Step 2
Multi-omics integration will connect expression with support layers
This follows preprocessing. It is not currently done.
M1 side
mRNA changes and support evidence for classical activation.
M2 side
mRNA changes and support evidence for alternative activation.
mRNA
- GSE162698
- GSE117040 validation
Methylation
- GSE117124
miRNA / prior art
- GSE51307
- GSE46903/GSE47189
Network, hub, regulator, and enrichment analysis are future steps
This is where regulator discovery would happen, but it has not happened yet.
Future analysis tools
After preprocessing, the project can use regulatory resources and network methods to rank transcription factors and other regulators. Known controls such as SPI1/PU.1, CEBPA/B, IRF8, and MAFB must be recovered before novel claims are trusted.
Network
- TF→target resources
- Regulon/activity scoring
- Hub detection
Enrichment
- GO/KEGG/Reactome
- M1/M2 pathway interpretation
Status
- Not started
- No final regulators
Final comparison: shared, M1-specific, and M2-specific regulators
This is the intended endpoint of Dr. SheikhBeig's workflow. It is still future work.
M1-specific candidates
Regulators stronger in M1 branch.
M2-specific candidates
Regulators stronger in M2 branch.
Shared regulators
Regulators active in both polarization paths, separated from branch-specific signals. This result has not been produced yet.
What the current page does and does not claim
This keeps the page scientifically honest for the professor.
Done
Database families were identified and confirmed internally. Dataset roles are proposed clearly on the page.
GreenWaiting for review by Dr. SheikhBeig
Step 2A–2D evidence package: metadata, matrices, donor mapping, and publication confirmation.
YellowNot claimed
No preprocessing, no differential expression, no M1/M2 network, no regulator discovery, no final candidate list.
FutureVisible records behind the proposed data stack
These are direct GEO, PubMed, or DOI links — not search-result links.