Waiting for review by Dr. SheikhBeig Databases found

Dr. SheikhBeig's workflow is now shown from A to Z.

The thesis is following the M1/M2 pathway Dr. SheikhBeig showed: monocyte/M0 → M1 and monocyte/M0 → M2 are treated as two branches, then compared. Right now we are still at the data-foundation stage: databases are found, dataset roles are proposed, and the evidence package is waiting for Dr. SheikhBeig's review.

👤 Mohammad Ali 🎓 Direct-entry PhD · Biotechnology 🕒 Page updated:

What pathway?

Dr. SheikhBeig's M1/M2 workflow

Two branches: M0/monocyte → M1 and M0/monocyte → M2, then compare.

What is done?

Database step is done

Confirmed internally and shown here. Relevant public resources were found.

Where are we now?

Data checks are waiting for review

Step 2A–2D checks are done internally and waiting for review by Dr. SheikhBeig.

What is not done?

No analysis yet

No preprocessing, no network, no regulator discovery, no final candidates.
First-look workflow picture

The whole workflow, read from top to bottom

One shared starting point splits into two parallel branches. Each branch runs down the same stages, and the two branches only come back together at the final candidate list. Status colors show where we are; the path itself does not change with progress.

Monocyte / M0

Shared starting point of both branches.

M1 branch

Monocyte/M0 → M1

Classical activation, usually IFNγ + LPS.

M2 branch

Monocyte/M0 → M2

Alternative activation, usually IL-4/IL-13 or IL-10.

Stage 1 · Omics & database gatheringDone
M1

Collect public records (GEO/SRA and support databases) holding M0 → M1 samples.

M2

Collect public records (GEO/SRA and support databases) holding M0 → M2 samples.

Stage 2 · Usable-data checksWaiting for review
M1

Check donors, sample labels, matrices, and papers behind the M1 arm.

M2

Check donors, sample labels, matrices, and papers behind the M2 arm.

Stage 3 · Preprocessing & QCFuture
M1

Reprocess raw reads, build count matrices, confirm M1 markers.

M2

Reprocess raw reads, build count matrices, confirm M2 markers.

Stage 4 · Integration & network buildingFuture
M1

Join mRNA with miRNA and methylation layers, then build the M1 network.

M2

Join mRNA with miRNA and methylation layers, then build the M2 network.

Stage 5 · Hub & regulator detectionFuture
M1

Rank M1 hubs and transcription-factor regulators.

M2

Rank M2 hubs and transcription-factor regulators.

Stage 6 · EnrichmentFuture
M1

GO/KEGG/Reactome meaning of the M1 modules.

M2

GO/KEGG/Reactome meaning of the M2 modules.

Both branches merge · M1 vs M2 comparisonFuture

Final candidate outputs

Shared regulators, M1-specific candidates, and M2-specific candidates, graded by evidence. This has not been produced yet.

Data & integration answers · Waiting for review by Dr. SheikhBeig

Data and integration checks requested by Dr. SheikhBeig

The points Dr. SheikhBeig raised are grouped into a short set of confirmation checks below, each with what was checked, what the result was, and the decision taken from it.

One-line summary: usable public M1/M2 macrophage datasets were found, but the workable design is evidence-layered multi-omics, not one fully matched multi-omics cohort. The strongest direct integration pair is GSE117040 (RNA-seq) with GSE117124 (methylation).
3 core omics layers Discovery backbone: GSE162698 Validation: GSE117040 Direct pair: GSE117040 + GSE117124 Fully matched cohort: no Regulator results: none yet
Grouped data and integration confirmation checks, results, and decisions
Confirmation checkResultDecision
Is the data actually available? Yes. GEO/NCBI, SRA, and PubMed were searched, with ENCODE, BLUEPRINT, and the Human Cell Atlas as reference resources. Usable public M1/M2 macrophage datasets were found, and each candidate has a defined role. The data-finding step is complete; work moves to how the datasets are used.
Which omics layers are usable? Three core layers: RNA-seq, DNA methylation (MBD-seq), and miRNA-seq. lncRNA may be extracted from RNA-seq where the annotation supports it. Report three core omics layers; describe lncRNA as RNA-seq-derived, not as a separate measured layer.
Are the datasets biologically and sample-wise compatible? Primary human macrophage, PBMC-derived macrophage, and THP-1 records were separated, and M0/M1/M2 structure was confirmed. GSE162698 is the strongest primary-human backbone (M0 plus M1 and M2 endpoints, donor structure, raw SRA data); GSE117040 is independent M1/M2 RNA-seq; GSE117124 shares its study framework and donor labels A–D. THP-1 lines, siRNA/transfected samples, and the unmatched miRNA cohort are not comparable to primary untreated human material. Use GSE162698 as the primary discovery backbone and GSE117040 as independent validation, kept separate rather than pooled. GSE117124 with GSE117040 is the compatible methylation pair. Non-comparable records stay support-only.
Can everything be integrated directly, and can batch effects be handled? No to full direct integration. Cell source, donor structure, perturbation, platform, and context differ, so the datasets are not one fully matched multi-omics cohort. Technical differences such as laboratory and platform can be modelled, but biological differences between non-comparable sources should not be corrected away. Use an evidence-layered integration design instead of claiming a matched cohort, and correct technical batch effects only where the underlying biology is comparable.
Where does mixOmics apply? Mainly to GSE117040 with GSE117124, where sample and design compatibility holds. It is weak for unrelated or unmatched datasets. Use mixOmics only for compatible blocks; use evidence-layered integration elsewhere.
What is the current boundary and the next gate? The stack covers general human macrophage M1/M2 polarization, not one shared disease-specific cohort. No network or regulator analysis has started. Make no regulator claims; this stage covers data and integration feasibility only. Disease-specific context stays a later extension unless compatible datasets are added.
Final data stack
Final data stack by role, dataset, and use
RoleDataset / resourceUse
Primary discovery backboneGSE162698Main M0 → M1 and M0 → M2 RNA-seq analysis.
Independent RNA-seq validationGSE117040Validate M1-vs-M2 signatures and candidate regulators.
Methylation support and direct integration candidateGSE117124 with GSE117040Best RNA + methylation integration candidate; possible mixOmics use.
miRNA supportGSE51307Post-transcriptional support layer, not a directly matched integration layer unless matching is shown.
Perturbation and context supportGSE329556Support and context only, because of the siRNA/transfection design and the absence of raw data.
Prior-art benchmarkGSE46903 / GSE47189Compare against the known macrophage activation literature.
Cell-line support onlyGSE273627 / GSE273628THP-1 support only; not used for primary-human claims.
Reference and backgroundENCODE / BLUEPRINT / HCA / CELLxGENEAnnotation and background only, unless a compatible matched dataset is selected.
Open point for Dr. SheikhBeig: this section is the data and integration answer as it currently stands. If a single disease-specific matched multi-omics setting is preferred instead, a separate targeted dataset search for that disease context would be needed.
Step 1 · Green dot · Done

Public databases were found and assigned to the workflow

This part is already confirmed internally. The point was to answer: “Which public databases can support Dr. SheikhBeig's pipeline?”

GEO / NCBI

Main source for human macrophage RNA-seq, miRNA-seq, methylation, ATAC/ChIP, and supplementary matrices.

Main dataset source

SRA / ENA

Raw sequencing reads behind GEO records. Needed when formal analysis must reprocess data rather than rely on TPM or normalized matrices.

Raw data source

ArrayExpress / BioStudies / OmicsDI

Cross-check source for older expression and multi-omics records not always visible through GEO-only search.

Cross-check

CELLxGENE / Human Cell Atlas

Single-cell/tissue atlas reference layer. Useful for projection/background, but not the controlled M1/M2 backbone.

Reference layer

ENCODE

Reference regulatory/chromatin data for monocytes and macrophages. Useful for annotation, not main M1/M2 contrast.

Annotation support

BLUEPRINT / EGA

Immune-cell epigenomic support layer; some resources are controlled access.

Epigenome support

PubMed / publisher records

Used to confirm each accession, paper, method, data availability, and prior art.

Paper confirmation
Output of this step: the database families are picked, the first dataset list exists, and the page now shows where each database fits into the workflow.
Step 2 · Blinking yellow dot · Waiting for review by Dr. SheikhBeig

Dataset checks are done internally; the proposed stack is ready for review

This step answers whether the found datasets are actually usable: human or cell line, donor structure, M0/M1/M2 labels, matrix files, raw data, and publication support.

2

What was checked?

Step 2A–2D checked GEO/GSM metadata, matrix availability, donor and column mapping, and publication-level confirmation.

Inputs checked

  • 647 GSM sample records
  • 13 GEO series records
  • 17 supplementary entries

Mapping checked

  • GSM sample titles
  • Matrix columns
  • Donor labels
  • M0/M1/M2 labels

Current decision

  • Done internally
  • Not approved by Dr. SheikhBeig yet
  • No analysis starts before review

Proposed dataset stack

These are the datasets picked into roles. They are proposed roles, not final locked analysis roles.

GSE162698

Primary discovery proposal. Primary human monocyte-derived macrophages; 3 donors; M0, M1, M2-IL4, M2-IL10, TAM-like. Processed TPM exists; raw SRA exists.

Discovery candidate

GSE117040

Independent RNA-seq validation. Donors A–D, paired M1/M2 RNA-seq. No M0 baseline, so it validates M1-vs-M2 rather than the full branch.

Validation

GSE117124

Methylation support. MBD-seq for the same donor-label structure as GSE117040; useful as an epigenetic support layer.

Methylation

GSE329556

Perturbation-context support. Primary human MDMs, siNC vs siXYLT2 across M0/M1/M2. Not pooled with untreated datasets; raw data unavailable.

Support only

GSE46903 / GSE47189

Prior-art benchmark. Xue et al. macrophage activation-spectrum resource. Important because it already contains network/regulator analysis.

Prior art

GSE51307

miRNA support. miRNA-seq arm with M0/M1/M2/TPP samples. Support layer only; not donor-matched to the discovery backbone.

miRNA

GSE273627 / GSE273628

THP-1 support only. These are cell-line records, not primary human donor datasets; they cannot carry discovery or validation claims.

Cell line only
Plain summary: databases are found, dataset roles are proposed, and the data evidence package is waiting for review by Dr. SheikhBeig. This still does not mean final datasets are locked or that analysis has started.
Step 3 · Grey dot · Future

Preprocessing starts only after Dr. SheikhBeig approves the data stack

This is the next practical step, but it has not started for the M1/M2 stack.

3

What will happen here?

Download/reprocess raw SRA where available, create count matrices, perform QC, confirm sample labels, and verify known M1/M2 marker behavior before network work.

Rule

  • Processed matrices are for preview/QC only
  • Formal DE/network work should use raw reprocessing where possible

Exception

  • GSE329556 has no raw data
  • Use only as support/context

Status

  • Not started
  • Waiting for Dr. SheikhBeig review of Step 2
Step 4 · Grey dot · Future

Multi-omics integration will connect expression with support layers

This follows preprocessing. It is not currently done.

M1 side

mRNA changes and support evidence for classical activation.

+

M2 side

mRNA changes and support evidence for alternative activation.

mRNA

  • GSE162698
  • GSE117040 validation

Methylation

  • GSE117124

miRNA / prior art

  • GSE51307
  • GSE46903/GSE47189
Step 5 · Grey dot · Future

Network, hub, regulator, and enrichment analysis are future steps

This is where regulator discovery would happen, but it has not happened yet.

5

Future analysis tools

After preprocessing, the project can use regulatory resources and network methods to rank transcription factors and other regulators. Known controls such as SPI1/PU.1, CEBPA/B, IRF8, and MAFB must be recovered before novel claims are trusted.

Network

  • TF→target resources
  • Regulon/activity scoring
  • Hub detection

Enrichment

  • GO/KEGG/Reactome
  • M1/M2 pathway interpretation

Status

  • Not started
  • No final regulators
Step 6 · Grey dot · Future

Final comparison: shared, M1-specific, and M2-specific regulators

This is the intended endpoint of Dr. SheikhBeig's workflow. It is still future work.

M1-specific candidates

Regulators stronger in M1 branch.

vs

M2-specific candidates

Regulators stronger in M2 branch.

Shared regulators

Regulators active in both polarization paths, separated from branch-specific signals. This result has not been produced yet.

Claim limits

What the current page does and does not claim

This keeps the page scientifically honest for the professor.

Done

Database families were identified and confirmed internally. Dataset roles are proposed clearly on the page.

Green

Waiting for review by Dr. SheikhBeig

Step 2A–2D evidence package: metadata, matrices, donor mapping, and publication confirmation.

Yellow

Not claimed

No preprocessing, no differential expression, no M1/M2 network, no regulator discovery, no final candidate list.

Future
Official sources

Visible records behind the proposed data stack

These are direct GEO, PubMed, or DOI links — not search-result links.