The current supervisor-facing focus is the M1/M2 database-first workflow: establishing whether public databases can support Dr. Sheikhbeig's monocyte/M0→M1 versus monocyte/M0→M2 pipeline.
That workflow — the public database search, the classification of each dataset by its role in the pipeline, the feasibility checks, and the two-branch analysis plan — is set out step by step on the main page. Start there.
This brief is a supporting document. It records the earlier pilot analyses on the human monocyte-to-macrophage axis and, more importantly, the evidence standards they established: which gates must pass before a regulator may be nominated, how curated-resource results are bounded, and how descriptive evidence is separated from causal evidence. Those standards carry forward to the M1/M2 pipeline.
Nothing in this brief is an M1/M2 result. No final regulators are claimed, no M1/M2 network has been built, and dataset roles for the M1/M2 pipeline are not locked. Donor and sample metadata verification remains the next gate.
Open the M1/M2 pipeline and database mapRead this first — what this page is, and is not
Status: pre-proposal / pilot feasibility evidence. This is not formal thesis completion. Everything shown here is preliminary computational work carried out to de-risk the study design and support a proposal. Formal thesis execution begins only after the proposal is approved and defended.
This page presents the recorded pilot feasibility evidence for the project. It contains no analysis, computation or result beyond what the recorded pilot analyses produced, and every number in the charts and tables is reproduced exactly as those analyses reported it. It is a supporting record; the current phase of work is the M1/M2 database-first workflow described in Current focus above and on the main page.
The downstream relevance option in §1B is a separate item put forward for the supervisory panel to evaluate; it is not the current phase of work and does not precede the M1/M2 database and feasibility phase. Every evidence ceiling and every gate stated on this page applies to it in full.
- that master regulators have already been found;
- that causal transcription factors have been proven;
- that initial monocyte→macrophage fate causality has been established;
- that any direct transcription-factor target has been identified;
- that any GRN edge or CellOracle simulation has been validated;
- that any candidate has been ranked or nominated;
- that therapeutic targets have been generated;
- that immunotherapy response — CAR-T or checkpoint-inhibitor — can be predicted;
- that the public clinical or perturbation resources named in §1B have been downloaded, processed, or had their metadata, patient labels or cell counts verified. Only their availability has been checked.
Those claims require gates that are not passed. The perturbation evidence below is descriptive and comes from already differentiated macrophages.
At a glance — the situation in one look
An executive summary of the recorded pilot evidence. The current phase of work is the M1/M2 database-first workflow; the items below are the methodological groundwork and evidence standards behind it, each stated in full further down the page.
Pilot evidence already established
Processed · replicatedTwo independent primary-human cohorts reproduce a cross-sectional monocyte→macrophage state ordering — Tier 1: 209,027 cells, 19 donors; Tier 2: 29,205 cells, 12 donors. Donor-blocked GRN diagnostics separate stable TF-level network context (out-degree Spearman ρ = 0.945 across disjoint donor folds) from unstable individual edges (Jaccard 0.126). This is a state-ordering model, not lineage tracing, real-time differentiation, or TF causality. §3 ↓ · §3B ↓
Key reliability result
Controls failedThe frozen curated CollecTRI/DoRothEA ULM/VIPER TF-activity layer failed the locked five-positive-control gate in both cohorts. No regulator nomination is allowed from that layer, anywhere in this project. It is kept as a documented, replicated result rather than repaired. §4 ↓
Proposed option for supervisor review
Proposed · not lockedRetain evidence-graded, donor-aware Mo→Mac regulatory modeling as the scientific core, and add a bounded projection of only tier-supported programs into immunotherapy-relevant macrophage programs. The clinical layer is a downstream relevance aim, never the project spine. §1B ↓
Why the option is feasible to discuss
Resources identified · availability verifiedPublic resources exist for each required layer: primary-human macrophage perturbation anchoring (Nat Genet 2024, linked accessions), TAM reference projection (Nat Commun 2024, linked Zenodo), CAR-T myeloid-resistance context (Cancer Cell 2025, linked Zenodo) and an ICI response cohort (GSE207422). Availability and eligibility only — none has been downloaded or processed here. §1B ↓
Current gate before this option can be locked
Metadata inspection pendingDownload-level metadata, patient/donor labels, cell counts, treatment/response labels and usable matrices must be inspected before the clinical option is locked; any resource failing those checks is dropped. The clinical output would be hypothesis prioritization — not response prediction and not target validation. A separate causal gate remains closed: the three primary-human perturbation benchmarks all act on already differentiated macrophages and establish no fate causality. §1B ↓ · §5 ↓ · §7 ↓
What feedback is requested now: the fifteen feedback prompts for supervisors ↓ — the first three, on whether the clinical projection is accepted as a bounded relevance layer, which macrophage domains to prespecify, and what metadata/cell-count threshold would be adequate to lock it, are the decisions to settle first.
1 · Proposed thesis scope and evidence boundaries
Core deliverable — donor-aware evidence framework
The valid, currently defensible core endpoint for human monocyte→macrophage regulatory inference:
Establish a reproducible, donor-aware human monocyte→macrophage state map across independent cohorts; test prospectively whether standard regulator-inference evidence layers recover known lineage regulators under locked positive controls; quantify what replicates, what fails, and what is only descriptive; and specify the wet-lab experiment required to resolve the remaining causal question.
This is the endpoint that today's evidence actually supports. It stands on its own whether or not the conditional extension beside it is ever opened.
Conditional extension — evidence-qualified candidate hypotheses
Permitted only if both conditions are met:
- A qualified direct human fate-transition perturbation resource — perturbation applied before or during monocyte→macrophage differentiation, with matched negative controls, ≥2 biological replicates (preferably donor matched), verifiable target efficacy, and full coverage of the fixed early and macrophage marker programs; and
- prospective causal controls — the positive-control rule written and frozen before any candidate ranking is inspected.
A professor-approved variant may instead accept primary-human macrophage-context perturbation as constrained causal-context support. If that variant is used, fate-causality language stays forbidden and must be replaced by explicit context wording.
Ceiling even if the gate opens Outputs are evidence-qualified candidate regulators — hypotheses carrying named supporting evidence and named limits. Never “master regulators”, never proven causal transcription factors, never therapeutic targets.
1B · Proposed clinical-projection option for supervisor review
Keep the evidence-graded Mo→Mac core, and add a bounded projection into immunotherapy-relevant macrophage programs
This section presents an option the panel is asked to evaluate. It is proposal-ready to discuss, and it is deliberately not presented as a settled part of the thesis: locking it depends on the resource-usability checks in the gate table at the end of this section.
The option in one sentence Build a donor-resolved, evidence-tiered evaluation of regulatory claims in the human monocyte→macrophage state axis, and project only the programs that survive that evaluation into prespecified, immunotherapy-relevant macrophage programs — as prioritized hypotheses for future validation.
Macrophage programs matter to immunotherapy
Cancer immunotherapy increasingly depends on, or is limited by, the state of tumour-associated myeloid and macrophage populations. Regulatory claims about macrophage reprogramming, however, are frequently derived from observational or model-specific evidence.
What this project would and would not say. It would ask which macrophage regulatory claims survive donor-aware stability, independent replication and perturbation-context anchoring before they are treated as clinically interesting. It would not claim new CAR-T or checkpoint-inhibitor biology, and it would not present itself as the first study of tumour-associated macrophages, of CAR-T resistance, or of macrophage reprogramming.
Donor-resolved, evidence-tiered regulatory claim evaluation
The spine of the thesis stays exactly what §1 describes: donor-aware modelling of the human Mo→Mac state axis, with each regulatory claim classified by the evidence that actually supports it — observational, network-predicted, donor-stable, independently replicated, perturbation-context supported, or directly causal.
“Donor-aware” has a tight definition here: the donor is the biological replicate; donor-blocked resampling and stability are computed; donor-level replication decides what may progress; and no pooling-first result may generate an unqualified claim.
Project only donor-stable, independently replicated, appropriately anchored programs
Nothing would be projected simply because it appeared in a network. A program would have to be donor-stable, independently replicated across cohorts, and appropriately anchored in perturbation context before it could enter the clinical layer at all. Projection targets would be prespecified before any clinical data are examined:
- Antigen presentation
- T-cell recruitment
- Suppressive / TAM-like programs
- Phagocytosis
- Inflammatory activation
The output is a prioritized hypothesis set The deliverable would be an uncertainty-annotated, prioritized set of macrophage regulatory hypotheses for future functional validation. It is not a CAR-T or ICI response predictor, not a validated therapeutic-target list, and not a claim of fate-causal master regulators.
D · Why the currently processed data support discussing this option
Two categories of evidence appear below and must not be confused. The left card is work already processed in this project. The right card is external, unprocessed — identified and checked for availability only.
- Reproducible state ordering in two independent primary-human cohorts (Tier 1: 209,027 cells / 19 donors; Tier 2: 29,205 cells / 12 donors) — §3B.
- Independent cohort design, with the donor as the resampling unit rather than the cell — §3A, §3D.
- A meaningful reliability failure retained rather than tuned away: the curated TF-activity layer failed the locked five-control gate in both cohorts — §4.
- Coarse GRN stability diagnostics from disjoint donor folds, separating stable TF-level ordering (ρ = 0.945) from unstable individual edges (Jaccard 0.126) — §3D.
- Primary-human post-differentiation perturbation benchmarks with verified target efficacy and complete fixed-program coverage — GSE155719, GSE224131, GSE181249 — §5.
Four public resources were identified as eligible for the projection layer on . What was verified is that each resource exists, is public, and carries the data types the option would need.
What has explicitly not been done: none of them has been downloaded or processed here, and no metadata field, patient/donor label, cell count, response label or matrix has been verified. That inspection is the next feasibility gate, not a completed step.
E · The four external resources, with stated use and stated limit
All four primary articles are peer-reviewed. Links open in a new tab; fuller citation context is in §2B, Group F.
Systematic perturbation screens of inflammatory macrophage states — Nature Genetics 2024Peer-reviewed
Use: perturbation anchoring of macrophage state — the layer that lets a regulatory claim be called perturbation-context supported rather than merely network-predicted.
Limit: mostly post-differentiation inflammatory-state context, not direct monocyte→macrophage fate causality. It can anchor macrophage state-maintenance and activation claims; it cannot license fate-causal language. One associated component is a controlled-access EGA study (EGAS00001006485) and is not assumed available.
DOI 10.1038/s41588-024-01962-w GSE210338 GSE210619 GSE210950 GSE268351 GSE268352 GSE268353 GSE268354 GSE270634 GitHub · Genentech/Haag_ng_2024 Zenodo 10.5281/zenodo.13836037 Zenodo source data 10.5281/zenodo.13836038
Coulton et al., Nature Communications 2024 — pan-cancer TAM atlasPeer-reviewed
Use: reference projection — mapping tier-passing programs onto an existing tumour-associated macrophage reference rather than building a new atlas.
Limit: the released atlas object carries no raw counts, so it supports signature/reference projection and not primary donor-aware modelling. No novelty claim over TAM biology and no response prediction may be built on it.
DOI 10.1038/s41467-024-49885-8 Zenodo · TAM Atlas 10.5281/zenodo.11222158 GitHub · alexcoulton/macrophage-atlas
Stahl et al., Cancer Cell 2025 — CSF1R+ myeloid-monocytic cells and CAR-T resistancePeer-reviewed
Use: disease-relevant myeloid / CAR-T context for the projection domain.
Limit: to be used as an external reference / projection resource only. This work already occupies the myeloid-driven CAR-T resistance space, so nothing here may be framed as discovering CAR-T resistance biology.
DOI 10.1016/j.ccell.2025.05.013 PMID 40513575 Zenodo 10.5281/zenodo.15280550
GSE207422 — NSCLC, pre/post PD-1 plus chemotherapyPeer-reviewed primary article
Use: a public NSCLC single-cell and bulk RNA resource from anti-PD-1 plus chemotherapy treatment, released with metadata — the candidate substrate for a bounded checkpoint-inhibitor projection.
Limit: it supports a bounded ICI projection only. It may not be used to build a response predictor; regulon-based myeloid response predictors already exist in the literature, and this project's endpoint is prioritization, not prediction.
Nothing in the clinical layer is committed until each row below is satisfied
| Item | Status | Required check before lock |
|---|---|---|
| Metadata | Not inspected | Open the released metadata for each candidate resource and confirm the fields the projection depends on are actually present and interpretable. |
| Patient / donor labels | Not inspected | Confirm patient or donor identity is recoverable per cell or per sample. Any resource without recoverable labels is excluded, because the donor is the biological replicate. |
| Myeloid / macrophage cell counts | Not quantified | Count myeloid and macrophage cells per patient and per condition, and check them against a threshold agreed with the panel before any projection is run. |
| Treatment / response labels | Not inspected | Confirm which treatment, timepoint and response annotations are released, and how complete they are across patients. |
| Usable matrix | Not inspected | Confirm an expression matrix in a usable form and state explicitly where raw counts are unavailable — as is already known for the TAM atlas object. |
| Response model | Excluded by rule | No response model is to be built. If response labels are sparse or incomplete, the affected projection is dropped rather than weakened — it is never converted into a predictor. |
- No claim of monocyte→macrophage fate causality, and no master-regulator nomination — the §4 control-gate failure and the §1 gate apply here in full.
- No claim of validated therapeutic targets, and no CAR-T or ICI response prediction.
- No primacy claim: not the first TAM atlas, not the first macrophage perturbation screen, not the first TAM/immunotherapy or CAR-T/myeloid link.
- Post-differentiation perturbation evidence does not become fate-causal evidence by being projected into a clinical domain.
- The clinical layer is a bounded downstream aim. If it were to become the primary endpoint, the work would collide with existing response-prediction literature — which is precisely why it stays downstream of the evidence-grading core.
2 · Positioning and claim limits
A targeted literature search found no exact duplicate of this core endpoint — and a great deal of adjacent work. The novelty claim therefore has to be made precisely, in both directions.
No primacy claims
- Not “the first macrophage gene-regulatory network”.
- Not “the first monocyte→macrophage atlas” or first transcriptomic analysis of this transition.
Human macrophage activation networks, polarization Boolean/logic models, human monocyte/macrophage atlases, generic single-cell GRN benchmarking, and SCENIC/GENIE3 myeloid-differentiation studies all exist. Claiming primacy over them would be false, and would be the easiest thing for a reviewer to reject.
The novelty is the combination, on one biologically specific axis
- human monocyte→macrophage state-axis focus;
- donor-aware two-cohort discovery/replication design;
- prospective positive-control gates for SPI1, CEBPA, CEBPB, IRF8 and MAFB, locked before inspection;
- explicit curated-resource failure analysis in the actual measured gene universe, rather than silent omission;
- donor-blocked GRN stability diagnostics separating stable coarse TF-level context from unstable individual edges;
- bounded primary-human perturbation context, integrated with its ceiling stated rather than implied;
- a proposed resolving wet-lab experiment for the question the computation cannot settle.
To our current knowledge this is a donor-aware reliability framework for human monocyte-to-macrophage regulatory inference. Unlike prior macrophage activation or polarization network studies, it tests prospectively whether standard TF-activity and GRN evidence can recover known lineage regulators under locked controls, evaluates replication across independent human donor cohorts, quantifies donor-blocked network stability, and separates descriptive, predictive and perturbational evidence before allowing any master-regulator claim.
Novelty risk. A reviewer may read this work as “just GRN benchmarking” or “just public-data reanalysis”. Mitigation: keep all five distinguishing features visible — biological specificity, donor-aware replication design, prospective control-gate failure reported as a result, bounded perturbation-context integration, and a concrete future wet-lab design. The work sits deliberately between generic GRN benchmarking (more biologically specific than that) and macrophage network-discovery papers (more methodologically strict than those).
2B · Why this positioning? Precedents and references
The reasoning behind the positioning
This subsection sets out why the scope boundaries in §1 are drawn where they are — the argument for the core deliverable, and for keeping the candidate-regulator extension behind a gate — so that the positioning can be argued rather than asserted. Every number in §3–§5 and every limit in §4, §5 and §6 applies exactly as stated there.
How the references were checked. Each item below was checked against its own published record before being cited here. The same standard applies to the clinical-resource group (§1B, Group F below): what was checked is that each resource is published and publicly released — not what its files contain.
It never needed to be first at anything
The surrounding literature is already occupied on both sides of this project:
- Reliability of regulatory inference is an established area. Reproducibility of single-cell GRN inference has been evaluated in peer-reviewed work since 2021, and modern GRN benchmarks exist — including a maintained “living” benchmark and a large benchmark built on single-cell perturbation data.
- Macrophage and monocyte regulatory biology is well characterised. Chromatin-defined human macrophage regulator networks, localized methylation change at TF binding sites across this transition, recent syntheses of human monocyte differentiation programmes, and MAFB as a conserved regulator of macrophage identity are all published.
The core endpoint is defensible precisely because it does not need to be first at any of these. It does not invent GRN reliability benchmarking, does not claim the first macrophage network, and does not discover MAFB. Its contribution is the stacked, transition-anchored, primary-human evidence architecture below — applied to one named human state transition, with its failures reported rather than repaired.
The collision is real, specific, and about wording
Recent literature already occupies MAFB/macrophage-identity biology, and a TAM regulatory-network preprint explicitly nominates MAFB as a key transcriptional driver in human tumour-associated macrophages. An extension framed as discovering MAFB, or as discovering TAM regulatory plasticity, would be both crowded and weak — independently of whether the §1 gate ever opens.
The consequence is a wording rule, not a scope change. The extension stays behind the closed §1 gate; if it opens, it must be framed around donor-aware, evidence-qualified regulator prioritisation in the human monocyte→macrophage transition — never around MAFB discovery, never around TAM plasticity discovery.
MAFB's role here is unchanged MAFB stays what it already is on this page: a locked positive control (§4) and a knockdown-benchmark target (§5). It is not a finding, in either case.
This is not merely “application novelty”
The core evidence framework is partly application/context novelty — but the substantive contribution is an evidence architecture that the adjacent literature does not stack together:
- Donor as the resampling unit. Reproducibility is measured across human donors — not across datasets, not across algorithm seeds — and cells from one donor are never treated as independent biological replicates.
- Independent cohort replication. The state axis and the activity layer are each re-tested in a separate human cohort (19 donors → 12 donors), so a finding must survive a cohort change, not just a re-run.
- Prospective lineage-control gates. SPI1, CEBPA, CEBPB, IRF8 and MAFB are locked as mandatory positive controls before inspection, so the pipeline can fail — and did.
- Curated-resource failure analysis. The control-gate failure is diagnosed in the actual measured gene universe (regulon overlap, cross-resource direction discordance) and retained as a result, not silently dropped or re-tuned.
- Donor-blocked GRN stability. Disjoint donor folds separate what is stable (coarse TF-level ordering) from what is not (individual edges), so network evidence is used only at the resolution it actually supports.
- Bounded perturbation-context interpretation. Primary-human perturbation benchmarks are admitted with their ceiling attached — post-differentiation, descriptive, no fate causality — rather than admitted and then quietly used as causal support.
Each element exists somewhere in the literature. The claim is the stack, anchored to one primary-human state transition, with each layer's permitted conclusion declared in advance.
Each reference below was checked against its own published record. For every item: what it is, why it matters here, and how it changes or limits what this thesis may claim.
Preprints are labelled with an orange Preprint tag and must not be cited as peer-reviewed; peer-reviewed items carry a green Peer-reviewed tag. All reference links open in a new tab.
Group ARegulatory-inference reliability and benchmarking — why this work may not claim to have invented the question
Kang, Thieffry & Cantini 2021 — Evaluating the reproducibility of single-cell gene regulatory network inference algorithms. Frontiers in Genetics 12:617282.Peer-reviewed
- Why it matters
- The closest published precedent for the reliability question asked here — a peer-reviewed evaluation of how reproducible single-cell GRN inference actually is.
- How it changes / limits the claim
- The thesis may not present “scGRN inference can be irreproducible” or “reproducibility should be measured” as its own insight, and must cite this work early. What stays distinct is the unit and the anchor: donors as the resampling unit inside one primary-human monocyte→macrophage transition, gated on prespecified lineage regulators — not algorithm-versus-algorithm reproducibility across datasets.
geneRNIB 2025 — A living benchmark for gene regulatory network inference.Preprint · bioRxiv
- Why it matters
- A maintained, community-facing GRN inference benchmark already exists, covering generic and context-specific settings.
- How it changes / limits the claim
- It forbids “the field lacks a GRN benchmark and this thesis supplies one”. This work is not a benchmark and must not be written as one; it is a transition-anchored reliability evaluation with biology-derived positive controls and donor-level replication. Cite as a preprint only.
CausalBench 2025 — A large-scale benchmark for network inference from single-cell perturbation data. Communications Biology 8:412.Peer-reviewed
- Why it matters
- Using perturbation data to evaluate network inference is established and has a large published benchmark behind it.
- How it changes / limits the claim
- No novelty may be claimed for “evaluating regulatory inference against perturbation evidence”. It also sets the honest contrast: CausalBench is large-scale cell-line perturbation used as benchmark ground truth, whereas §5 here is three small primary-human, post-differentiation datasets used as bounded context — a weaker causal position that must be stated as such, not dressed up as a benchmark.
Müller-Dott et al. 2023 — CollecTRI TF regulons. Nucleic Acids Research.Peer-reviewed
- Why it matters
- This is the source resource for the TF-activity layer that failed the five-control gate in both cohorts (§4). Its authors already evaluated regulon quality against perturbation data.
- How it changes / limits the claim
- §4 must be stated as a context-specific reliability finding in this transition — not as a refutation of CollecTRI, and not as a general verdict on TF-activity estimation. Conversely, it is what makes §4 reportable rather than a bug report: the failure was found under prospectively locked controls, replicated in an independent human cohort, and diagnosed as regulon overlap and cross-resource direction discordance in the measured gene universe.
Group BMonocyte→macrophage and macrophage regulatory biology — why no primacy claim is available
Dekkers et al. 2019 — Human monocyte-to-macrophage differentiation involves highly localized gain and loss of DNA methylation at transcription factor binding sites. Epigenetics & Chromatin.Peer-reviewed
- Why it matters
- The epigenomic layer of this exact transition is already characterised at TF-binding-site resolution. This study is also the source of Tier 1 dataset GSE118696.
- How it changes / limits the claim
- The regulatory/epigenomic landscape of monocyte→macrophage differentiation may not be presented as uncharted. The temporal chromatin anchor stays exactly what §6 already says it is — descriptive, with no differential-accessibility and no causal claim.
Schmidt et al. 2016 — The transcriptional regulator network of human inflammatory macrophages is defined by open chromatin. Cell Research 26(2):151–170.Peer-reviewed
- Why it matters
- A chromatin-anchored human macrophage transcriptional-regulator network already exists, built from matched RNA-seq and histone ChIP-seq.
- How it changes / limits the claim
- It is the direct reason for the “not the first macrophage gene-regulatory network” prohibition in §2, and it means constructing a chromatin-informed regulator network is not itself a contribution. The distinction to state is scope and standard of evidence: activation-condition network construction there, versus donor-aware reliability evaluation across the monocyte→macrophage state axis here.
- Citation note
- The correct lead author is Schmidt SV; this paper is mis-attributed elsewhere. Cite it as Schmidt et al. 2016.
Komaravolu et al. 2025 — Transcriptional programs underlying human monocyte differentiation and diversity. Journal of Leukocyte Biology.Peer-reviewed
- Why it matters
- The closest recent overlap on human monocyte differentiation programmes and their transcriptional regulators, including SCENIC-style regulator analysis.
- How it changes / limits the claim
- The control TFs used here are prior knowledge and must be presented that way. Recovering SPI1/CEBPA/CEBPB/IRF8/MAFB is a pipeline validity check, never a finding — which is exactly why their non-recovery in §4 is informative and their recovery would not have been.
Vanneste et al. 2026 — MafB is a conserved transcriptional regulator of macrophage development and functional identity across tissues and species. Immunity.Peer-reviewed
- Why it matters
- MAFB's role in macrophage identity is established across tissues and species.
- How it changes / limits the claim
- A hard stop on any MAFB-discovery framing, in the core deliverable and in the conditional extension alike. MAFB stays a locked positive control (§4) and a knockdown-benchmark target (§5.1, §5.2); the GSE155719 and GSE224131 values remain descriptive post-differentiation knockdown-efficacy context and gain no causal weight from this reference.
Group CIn-silico perturbation method and its caution — why CellOracle stays gated
CellOracle 2023 — Dissecting cell identity via network inference and in silico gene perturbation. Nature.Peer-reviewed
- Why it matters
- The published method the gated final stage in §6 would use.
- How it changes / limits the claim
- Running a published method is not novelty. It keeps CellOracle inside its gate — permitted only after the conditional-extension evidence rule passes, or as a clearly labelled methodology demonstration carrying no nomination, no ranking, and no fate-shift claim.
Critique of CellOracle — Critical issues found in “Dissecting cell identity via network inference and in silico gene perturbation”.Preprint · bioRxiv · caution only, not consensus
- Why it matters
- It raises specific methodological concerns about the original CellOracle analyses, and supports the decision to gate rather than run.
- How it changes / limits the claim
- It is an unrefereed preprint. It may be cited as caution and as part of the rationale for gating; it may not be cited as a settled refutation, as community consensus, or as evidence that CellOracle results are invalid.
Group DCollision context — the specific paper that constrains how the extension may be worded
Regulatory networks shaping human tumor-associated macrophages in vivo identify MAFB as a key transcriptional driver.Preprint · Research Square · read in full before formal citation
- Why it matters
- The specific, named collision for a TAM/MAFB-centred extension: it nominates MAFB as a key transcriptional driver in human tumour-associated macrophages using regulatory-network evidence.
- How it changes / limits the claim
- The conditional extension must not be framed as discovering MAFB or TAM regulatory plasticity. It does not invalidate this project: it appears TAM/in-vivo centred rather than a donor-aware reliability framework for the monocyte→macrophage transition. It is a preprint and must be read in full before it is cited formally or used to argue either overlap or distinction.
Group EEvidence datasets cited in §5 — verify the designs that set the ceilings
Linked so a reader can check the design features — donor labelling, differentiation timing, released data type — that set the interpretation ceilings in §5. All three perturb already differentiated macrophages; nothing in this section relaxes that ceiling.
GSE155719 — MAFB siRNA, three donor-labelled pairs, released read counts
- Why it matters
- The strongest primary-human perturbation benchmark on this page (§5.1): explicit donor labels and count tables, which is what makes the donor-paired log2 ratios formable inside a donor.
- How it changes / limits the claim
- The record shows seven days of M-CSF differentiation before siRNA — the reason the ceiling says post-differentiation, descriptive, no initial fate causality.
GSE224131 — MAFB siRNA, GM-CSF and M-CSF macrophage contexts
- Why it matters
- Provides a second, two-cytokine-context MAFB benchmark (§5.2).
- How it changes / limits the claim
- The released files encode
rep1–3without donor IDs, which is why those values are labelled replicate-index contrasts rather than verified donor-paired effects.
GSE181249 — AHR siRNA, donor-matched, TLR7 activation context
- Why it matters
- Shows a donor-matched primary-human perturbation design with a full prespecified-gene universe (§5.3).
- How it changes / limits the claim
- It is an activation experiment on already differentiated M-CSF macrophages, so it supports activation-context interpretation only — not fate causality, not direct targets, not ranking.
Group FClinical and perturbation resources for the proposed projection option — peer-reviewed, availability verified 2026-08-13
These four peer-reviewed resources are what make the §1B option discussable. Each was confirmed to be published and publicly released on . None has been downloaded or processed, and no metadata field, patient label, cell count or matrix has been verified — that is the gate in §1B·F, not a completed step.
Haag et al. 2024 — Systematic perturbation screens identify regulators of inflammatory macrophage states and a role for TNF mRNA m6A modification. Nature Genetics.Peer-reviewed
- Why it matters
- It is the strongest available basis for a real perturbation-anchoring layer in primary human macrophages, with released GEO accessions, analysis code and Zenodo archives — which is what would let a claim be graded “perturbation-context supported” rather than only network-predicted.
- How it changes / limits the claim
- Use: perturbation anchoring of macrophage state. Limit: mostly post-differentiation inflammatory-state context, not direct Mo→Mac fate causality; it narrows novelty rather than creating it, so no “first macrophage perturbation atlas” framing is available. One associated component is controlled-access EGA (
EGAS00001006485) and is not assumed available.
doi.org/10.1038/s41588-024-01962-w GSE210338 GSE210619 GSE210950 GSE268351 GSE268352 GSE268353 GSE268354 GSE270634 github.com/Genentech/Haag_ng_2024 zenodo.13836037 zenodo.13836038 · source data
Coulton et al. 2024 — Using a pan-cancer atlas to investigate tumour associated macrophages as regulators of immunotherapy response. Nature Communications.Peer-reviewed
- Why it matters
- A public TAM atlas with an accompanying code repository already exists, so the projection domain does not need a new atlas built for it.
- How it changes / limits the claim
- Use: reference projection. Limit: the released atlas object contains no raw counts, so it cannot serve as primary donor-aware modelling input; and because TAM/immunotherapy-response association is already occupied by this work, projection must stay bounded as hypothesis prioritization with no novelty or response-prediction claim.
doi.org/10.1038/s41467-024-49885-8 zenodo.11222158 · TAM Atlas github.com/alexcoulton/macrophage-atlas
Stahl et al. 2025 — CSF1R+ myeloid-monocytic cells drive CAR-T cell resistance in aggressive B cell lymphoma. Cancer Cell. PMID 40513575.Peer-reviewed
- Why it matters
- It supplies public single-cell, bulk and protein data for a myeloid-driven CAR-T resistance setting — the most directly immunotherapy-relevant myeloid context among the four.
- How it changes / limits the claim
- Use: disease-relevant myeloid / CAR-T context. Limit: external reference and projection only. It is also the clearest prior-work collision in this space, so no discovery of CAR-T resistance biology may be claimed.
doi.org/10.1016/j.ccell.2025.05.013 pubmed.ncbi.nlm.nih.gov/40513575 zenodo.15280550
GSE207422 — NSCLC anti-PD-1 plus chemotherapy single-cell and bulk RNA resource; primary article in Genome Medicine 2023.Peer-reviewed primary article
- Why it matters
- A public checkpoint-inhibitor cohort with a released single-cell UMI matrix and metadata plus bulk expression and metadata — the candidate substrate for a bounded ICI projection.
- How it changes / limits the claim
- Use: bounded ICI projection of tier-passing programs. Limit: not a response predictor. Regulon-based myeloid predictors of immunotherapy response are already published, so a predictive endpoint here would be both crowded and outside this project's permitted claim scope.
- These precedents are cited early in the proposal, not buried in a discussion section.
- The novelty sentence in §2 stands — with the six-part stack above named explicitly as the contribution.
- The conditional extension's wording is constrained as above, independently of whether the §1 gate opens.
- The clinical resources in Group F are cited as identified public resources supporting the §1B option — never as data this project has already analysed.
- No evidence value, interpretation ceiling, or gate in §3–§6 changes.
3 · What is already demonstrated (pilot feasibility)
A well-powered human single-cell discovery resource exists
- Tier 1 discovery: 209,027 QC-eligible singlet cells, 19 donors, 198 libraries, classical monocytes, non-classical monocytes, and macrophages.
- Tier 2 replication: 29,205 locked primary-lineage cells from 12 independent donors, with all three states represented for every donor.
- Raw data are immutable; processed artifacts, manifests, checksums, cell-accounting, and memory-safe workflows are in place.
The biological state axis is reproducible across independent cohorts
Both tiers independently reproduce the expected cross-sectional state ordering. This validates a reproducible state-ordering model — not lineage tracing, real-time differentiation, or TF causality.
Orthogonal molecular layers are available
- Temporal bulk RNA (0–144 h) shows expected early-monocyte decline and macrophage-marker increase; TF expression there is orthogonal and descriptive only.
- Temporal ATAC provides a descriptive chromatin anchor (n=1 track per timepoint), with strict no-causality / no-differential-accessibility limitations.
- A full donor-aware GRNBoost2 network is technically validated on 4,280 metacells × 5,528 genes. Its importances are unsigned predictive associations — not binding, not activation/repression, not causality.
- Perturbation data span HL-60 model-system positive controls and three primary-human post-differentiation benchmarks — see §5 ↓.
Reproducibility infrastructure is already demonstrated
- Library-only Harmony integration was validated, rather than correcting biological labels.
- Donor-level endpoint pseudobulk design avoids treating cells from the same donor as independent biological replicates.
- Donor-blocked GRN diagnostics quantify what is stable (coarse TF-level ordering; out-degree Spearman ρ = 0.945 across disjoint donor folds) and what is not (individual inferred edges; Jaccard 0.126, shared-edge importance ρ = 0.285).
- All major pilot outputs have manifests, explicit input checksums, and raw-data immutability checks.
3B · Reproducible state ordering — Tier 1 vs Tier 2
Median DPT and the ordering probability are different measures that happen to share the 0–1 range. Compare the two tiers within a metric; do not compare values across metrics.
| Validation | Tier 1 | Tier 2 |
|---|---|---|
| Classical-monocyte median DPT | 0.020 | 0.068 |
| Non-classical-monocyte median DPT | 0.122 | 0.256 |
| Macrophage median DPT | 0.189 | 0.345 |
| P(macrophage orders later than classical monocyte) | 0.9917 | 0.9469 |
| Canonical early/late marker trends | passed | passed |
Interpretation ceiling. This validates a reproducible cross-sectional state-ordering model. It is not lineage tracing, not observed real-time differentiation, not differentiation rate, and not evidence of TF causality. This limit is unchanged from the 2026-08-09 brief.
4 · The important negative / reliability result — preserved and strengthened
The frozen CollecTRI/DoRothEA ULM/VIPER TF-activity layer failed the prespecified five-positive-control gate in both independent human cohorts
CEBPA was rank 175/789 (top 22.2%), outside the locked top-20% rule. SPI1, CEBPB, IRF8 and MAFB fell inside the top 20% but showed cross-resource direction discordance and failed the two-independent-source evidence rule.
Among 777 ranked TFs, only SPI1 passed (rank 41). CEBPA (456), CEBPB (298) and IRF8 (188) failed the rank rule; MAFB (99) failed the direction-concordance rule.
Diagnosis. Methods agree within a resource (Spearman 0.73–0.83) but not across resources (0.34), driven by low control-regulon overlap (Jaccard 0.010–0.116; MAFB shares a single target) — not an obvious scoring bug.
- No threshold, resource, method, gene universe or ranking rule may be re-tuned to make the controls pass retrospectively.
- The failed layer may not be combined post hoc — by weighting, ensembling, rank aggregation, or a composite multi-evidence score — into anything presented as passing evidence.
- No master regulator may be nominated from this layer, in either cohort, anywhere in this project.
- The one permitted use: a documented, replicated reliability finding that generic curated regulon activity cannot serve as a stand-alone master-regulator discovery engine in this biological setting.
This is not a reason to abandon the project. It is why the design is now evidence-gated: independent evidence layers must support a candidate before it is promoted, and the failed layer never counts among them.
5 · Primary-human perturbation evidence — descriptive, post-differentiation
Three primary-human perturbation resources now pass technical, coverage and target-efficacy checks
All three perturb already differentiated macrophages. They are context / positive-control benchmarks. None of them tests the initial monocyte→macrophage fate switch, and none of them opens the conditional extension.
Bars run from a zero baseline: left of zero is lower in siMAFB than in that donor's own siControl, right of zero is higher. Every value is a descriptive log2 ratio — no model was fitted, no test was run, and no p-value, FDR or significance claim exists.
| Measure | Donor #1 | Donor #2 | Donor #3 | Mean |
|---|---|---|---|---|
| MAFB — perturbation target | −2.176 | −2.093 | −1.893 | −2.054 |
| Early-monocyte program (5/5 genes) | +0.525 | +0.648 | +0.551 | +0.574 |
| Macrophage program (6/6 genes) | −0.260 | −0.273 | −0.267 | −0.267 |
| SPI1 — contextual reference, in no program | +0.011 | +0.073 | −0.016 | +0.023 |
Primary human CD14+ monocytes from three donors, differentiated seven days with M-CSF, then transfected for 24 h with control or MAFB-specific siRNA. Six released HTSeq read-count tables, three donor-labelled siControl/siMAFB pairs, one identical 58,676-accession Ensembl universe, 13/13 contract-frozen genes present in every table. The MAF-specific siRNA arm was excluded.
PASS 3/3 donors — MAFB log2 ratio negative in all three (range 0.283). Raw MAFB counts fell 17,187→3,814, 16,925→4,174 and 14,707→3,938. The early program moved up and the macrophage program moved down in all three donors.
- Every value is descriptive. n = 3 donor pairs; no DESeq2/edgeR/limma model was fitted, no test was run, no p-value, FDR or significance claim exists anywhere.
- The perturbation is post-differentiation (7 days M-CSF before siRNA), so no initial fate causality can be read from it — no fate transition occurs inside the experiment's window, and a program moving the way a differentiation model would predict is not a confirmation of that model here.
- No direct MAFB target is established; an expression change after knockdown is consistent with many indirect routes.
- No GRN or CellOracle validation. No GRNBoost2 edge and no simulation is validated, and this does not rescue the Tier 1 / Tier 2 curated-resource failure.
- No candidate ranking or nomination, and no therapeutic claim. All thirteen genes and both programs were fixed by contract before any count was read.
- CPM inherits its own assumptions (library size = ENSG-row total), and the excluded MAF arm means nothing here speaks to MAF.
Primary-human GM-CSF and M-CSF macrophage contexts
- Uninfected subset only; three matched contrasts per cytokine context; 17,050-gene universe; 13/13 prespecified symbols resolved.
- MAFB suppression negative in all three contrasts in both contexts — GM-CSF −0.510, −0.785, −0.612 (mean −0.635); M-CSF −0.567, −0.893, −0.686 (mean −0.715).
- Program means: GM-CSF early −0.144, macrophage −0.124; M-CSF early +0.681, macrophage −0.169. SPI1 reported separately (+0.045, −0.051).
rep1–3 without donor IDs, so these are replicate-index contrasts, not verified donor-paired effects; perturbation follows six days of differentiation. Descriptive FPKM ratios only.Primary-human M-CSF macrophages under TLR7 activation
- Three donor-matched siControl/siAHR pairs at baseline and after the TLR7 ligand CL264; 17,427-gene universe; 13/13 required genes covered.
- AHR suppression met the prespecified direction-only rule in 3/3 donors at baseline (mean −1.882 log2 FPKM) and under CL264 (mean −2.621).
- Program means: baseline early −0.381, macrophage −0.078; CL264 early +0.120, macrophage +0.069. Fixed CL264-minus-baseline interaction: early +0.501 (3/3 donors), macrophage +0.147 (mixed).
What this evidence does and does not buy
It does demonstrate that a controlled primary-human perturbation layer can be integrated into this thesis without overclaiming: matched designs, verifiable target efficacy, and complete fixed-program coverage are all achievable on public data.
It does not open the conditional extension. Direct perturbation during human monocyte→macrophage differentiation remains unfound and unproven.
6 · Evidence-tier architecture — tier selector
Select a tier to see its purpose, its current status, and — critically — the conclusion it is permitted to support. Each tier's ceiling is set by the brief, not by convenience.
Tier 1 — Human donor-aware discovery
Data / QC / trajectory / GRN completed
Human donor-aware discovery.
209,027 QC-eligible singlet cells · 19 donors · 198 libraries · classical monocytes, non-classical monocytes, and macrophages.
- State axis and coarse network context.
- Lineage tracing.
- Real-time differentiation.
- TF causality.
Tier 2 — Independent human validation
Trajectory completed; activity layer replicated its limitation
Independent human validation.
29,205 locked primary-lineage cells · 12 independent donors · all three states represented for every donor.
- Independent state-axis validation.
- Negative reliability result.
- Lineage tracing.
- Real-time differentiation.
- TF causality.
Tier 3 — Temporal, chromatin, perturbation, disease/TAM extensions
Temporal anchors completed; three primary-human perturbation benchmarks completed
Temporal, chromatin, perturbation, and disease/TAM extensions.
Temporal bulk RNA (0–144 h); temporal ATAC as a descriptive chromatin anchor (n=1 track per timepoint); HL-60 model-system positive controls; and three primary-human post-differentiation perturbation benchmarks — GSE155719 (MAFB, donor-paired, count-based), GSE224131 (MAFB), GSE181249 (AHR). See §5 ↑
- Contextual / descriptive support only, unless prospective gates pass.
- Perturbation-efficacy and macrophage-context marker behaviour, explicitly labelled post-differentiation.
- Causality or differential-accessibility claims from the ATAC layer.
- Initial monocyte→macrophage fate causality from any perturbation benchmark.
- Direct TF targets, candidate ranking, or therapeutic claims.
Clinical projection layer — proposed option, awaiting supervisor review
Proposed · not locked Resource metadata not inspected
Bounded projection of tier-supported regulatory programs into prespecified immunotherapy-relevant macrophage programs — antigen presentation, T-cell recruitment, suppressive/TAM-like programs, phagocytosis, and inflammatory activation. See §1B ↑
Primary-human macrophage perturbation anchoring (Nat Genet 2024, with linked GEO/GitHub/Zenodo releases); TAM reference projection (Nat Commun 2024 atlas on Zenodo); CAR-T myeloid-resistance context (Cancer Cell 2025 on Zenodo); and an ICI response cohort (GSE207422). None has been downloaded, processed, or had its metadata, labels or cell counts verified.
A program enters this layer only if it is donor-stable, independently replicated, and appropriately anchored. The layer itself is only locked once the §1B·F gate passes — metadata, patient/donor labels, myeloid/macrophage counts, treatment/response labels, and a usable matrix.
- Prioritized, uncertainty-annotated macrophage regulatory hypotheses.
- Statements of clinical relevance for programs that already passed the earlier tiers.
- CAR-T or checkpoint-inhibitor response prediction.
- Validated therapeutic targets.
- Any primacy claim over TAM, CAR-T/myeloid or macrophage-perturbation literature.
- Projection of programs that did not pass the earlier tiers.
Final causal stage — Direct-fate perturbation evidence + CellOracle
Gated / not started
Direct human fate-transition perturbation evidence combined with CellOracle in-silico perturbation.
This stage does not begin because the calendar reaches it. It begins only when the conditional-extension gate is passed: a qualified direct human fate-transition perturbation resource and a prospective causal-control rule locked before any candidate ranking is inspected.
Running it now would produce simulations with no qualified causal anchor. It may be used only (a) after the prospective evidence rule passes, or (b) as a clearly labelled methodology demonstration inside the core framework — carrying no candidate nomination, no ranking, and no fate-shift claim.
- Nothing until the gate opens.
- Then: candidate fate-shift predictions for evidence-qualified candidates only.
- Any candidate nomination if prospective controls fail (explicit stop rule).
- A manufactured master-regulator list.
- Unlabelled CellOracle output presented as validation.
7 · What remains before any final regulator claim
- Deliver the core evidence framework in full — the reliability framework, its replicated control-gate failure, its donor-blocked stability diagnostics, its bounded perturbation-context benchmarks, and its evidence rules. This is the endpoint; it is not contingent on anything below.
- Keep scouting for a direct human fate-transition perturbation resource — perturbation before/during differentiation — without letting proposal feasibility depend on finding one.
- Lock one prospective evidence-synthesis rule before any candidate ranking is inspected, separating: Tier 1 donor-aware endpoint behaviour; Tier 2 independent behaviour; coarse TF-level GRN context; trajectory/temporal/chromatin as descriptive support; and direct perturbation as the causal-context layer. The failed curated layer cannot enter it.
- Run the clinical-resource usability inspection before the §1B option is locked — metadata, patient/donor labels, myeloid/macrophage counts per patient and condition, treatment/response labels, and a usable matrix for each candidate resource. Resources that fail are dropped; no response model is built in any case.
- Only if the conditional-extension gate passes: constrained candidate qualification, then gated CellOracle in-silico perturbation for eligible candidates, then disease/TAM and therapeutic-knowledge-base context — all described as evidence-qualified hypotheses.
- Specify the resolving wet-lab experiment in either case: primary human monocytes perturbed before/during M-CSF differentiation, with readouts for target suppression, state-marker programs, and single-cell state shift. If the §1B option is adopted, one prespecified immunotherapy-relevant macrophage domain — antigen presentation, T-cell recruitment, suppressive/TAM-like, phagocytosis, or inflammatory activation — is added as a readout.
Stop rule If the causal gate never opens, the thesis stands as the core evidence framework — a rigorous donor-aware multi-omics reliability study with a documented, replicated limitation of curated TF-activity models. It must not manufacture a master-regulator list. Independently: if the clinical-resource usability gate fails, the §1B projection layer is dropped rather than reduced to a weaker version of itself, and the core endpoint is unaffected.
8 · Proposal-ready objectives
Obj 1Build the human monocyte-to-macrophage reference state map
Integrate Tier 1 and independent Tier 2 public human single-cell cohorts using donor-aware QC, library-only batch correction, and state-ordering validation.
Obj 2Reconstruct and evaluate regulatory-network evidence
Infer a data-driven TF–target association network from donor-aware metacells; quantify donor-blocked stability; assess curated TF-activity resources against mandatory positive controls.
Obj 3Establish orthogonal evidence for candidate regulators
Integrate temporal RNA/chromatin anchors and qualified direct perturbation datasets under a prospective positive-control framework.
Obj 4Predict and prioritize experimentally testable consequences
For candidates passing Objective 3, perform in-silico perturbation, disease/TAM extension, therapeutic-context mapping, and specify one primary-monocyte validation experiment.
Obj 5 · proposedBounded clinical projection into immunotherapy-relevant macrophage programs
This objective is the option in §1B and is offered for the panel to accept, reshape or decline. It would project only donor-stable, independently replicated, appropriately anchored programs onto the five prespecified macrophage domains, using public TAM, CAR-T/myeloid and ICI resources together with primary-human macrophage perturbation anchoring.
It is a bounded downstream aim, never the project spine, and it does not begin until the resource-usability gate in §1B·F passes.
9 · Feasibility and risk matrix
Nine dimensions, each with the brief's own assessment label and the mitigation or evidence that backs it. Filter by assessment.
Data availability
StrongTwo large independent human scRNA cohorts already acquired; temporal RNA/ATAC and perturbation extensions available.
Computational feasibility
StrongWorkflows have run successfully at full Tier 1 scale with bounded memory and reproducible publication.
Biological state-map validity
StrongIndependent state ordering and canonical marker trends reproduced.
Direct TF-causality evidence
Current bottleneckCurated activity models failed controls; qualify external perturbation data before candidate claims.
Multi-omics integration depth
ModerateRNA/trajectory and chromatin anchors exist; paired primary-human multiome would strengthen but is not prerequisite for a viable hybrid thesis.
Final candidate nomination
Not yet establishedExplicit stop rule: no nomination if prospective controls fail.
Public clinical-resource availability
ModeratePeer-reviewed, publicly released resources were identified for each layer the §1B option needs — macrophage perturbation anchoring, TAM reference projection, CAR-T myeloid context, and an ICI cohort. Availability and eligibility only; nothing downloaded or processed.
Clinical-resource metadata usability
Current gate for the proposed optionPatient/donor labels, myeloid-macrophage counts, response labels and matrix usability are all uninspected. Mitigation: the §1B·F gate must pass before the option is locked, and any resource failing it is dropped.
Clinical / therapeutic extension
Future conditional workBounded projection is proposed as a downstream relevance layer, not a pilot claim, and yields prioritized hypotheses only — never response prediction or validated targets. Candidate-regulator work still begins only after candidate qualification.
No dimension carries that assessment.
10 · Recommended milestone gates after proposal approval
These gates begin after the proposal is approved. None of them is claimed as complete.
-
Gate A — Data / reference lock
- Success criterion
- All source manifests and Tier 1 / Tier 2 scope frozen.
- Decision
- Begin formal analysis
-
Gate B — State-map reproducibility
- Success criterion
- Independent marker / order checks retained.
- Decision
- Proceed to regulator evidence
-
Gate C — Core evidence-framework delivery
- Success criterion
- Control-gate outcomes, donor-blocked GRN stability, bounded perturbation-context benchmarks and the evidence-tier rules all documented.
- Decision
- Core endpoint secured, independently of everything below
-
Gate D — Clinical-resource usability proposed option
- Success criterion
- For each candidate resource: metadata inspected, patient/donor labels recoverable, myeloid/macrophage counts per patient and condition above the panel-agreed threshold, treatment/response labels characterised, and a usable matrix confirmed.
- Decision
- Lock the §1B projection layer, or drop the failing resource
-
Gate E — Causal-evidence qualification conditional
- Success criterion
- A qualified direct human fate-transition perturbation resource and a prospective causal-control rule locked before candidate inspection.
- Decision
- Open the conditional extension; otherwise deliver the core framework endpoint
-
Gate F — CellOracle gate
- Success criterion
- Gate E passed and the predeclared evidence rule passed.
- Decision
- Permit in-silico perturbation for eligible candidates; otherwise a labelled methodology demonstration only
-
Gate G — Final reporting
- Success criterion
- Evidence-qualified candidate set or reliability/framework result, plus the specified primary-monocyte experiment; plus prioritized macrophage-program hypotheses if the §1B option was locked at Gate D.
- Decision
- Thesis / manuscript and wet-lab proposal
11 · Honest proposal status statement
The wording below is the brief's own status statement, intended for use verbatim in the proposal.
For the proposal document Preliminary computational feasibility work has established access to large independent human single-cell cohorts, reproducible monocyte–macrophage state ordering, temporal RNA/chromatin context, a donor-aware network-analysis pipeline, and three primary-human perturbation benchmarks including donor-paired MAFB knockdown with raw counts. The pilot also identified a key methodological limitation: two curated TF-activity resources did not recover all prespecified positive controls across independent cohorts, and that failure is retained rather than tuned away. The available perturbation resources act on already differentiated macrophages, so they support context interpretation only. The proposed thesis is therefore a donor-aware, evidence-graded framework for regulatory inference in this state transition, in which regulator nomination requires independent causal and replication support rather than a single resource-derived score. A bounded clinical projection of tier-supported programs onto prespecified immunotherapy-relevant macrophage programs is proposed as a downstream relevance layer, using public macrophage-perturbation, tumour-associated-macrophage, CAR-T myeloid and checkpoint-inhibitor resources whose availability has been verified but whose metadata, patient labels, cell counts and matrices remain to be inspected; its output would be prioritized hypotheses for future validation, not immunotherapy-response prediction or validated therapeutic targets.
Meeting aid · Feedback prompts for supervisors
Fifteen questions to put to the supervisory panel, each tied to a specific decision the brief leaves open. The first three concern the proposed clinical-projection option in §1B and are the decisions to settle first. Expand a prompt to see why it matters and what the brief already says.
Nothing on this page is submitted, transmitted, or stored anywhere. These are talking points only — capture answers in your own notes.
Clinical optionDo you accept the proposed clinical projection as a bounded downstream relevance layer on top of the evidence-graded core — or should the thesis stay entirely within the monocyte→macrophage evidence framework?
This is the first decision, because everything else in §1B depends on it. Accepting the option adds one conditional objective and one milestone gate; declining it removes both, and the core endpoint is unaffected either way.
The specific thing to test: is a projection whose output is a prioritized hypothesis set — not a response predictor, not a target list — worth the added scope, given that the clinical resources have not yet been inspected?
Projection domainsWhich immunotherapy-relevant macrophage domains should be prespecified — and should any be added or removed before any clinical data are examined?
Five domains are currently proposed: antigen presentation, T-cell recruitment, suppressive / TAM-like programs, phagocytosis, and inflammatory activation. Fixing the list in the proposal meeting is what makes the projection prespecified rather than chosen once the data have been seen.
A narrower set would be a strengthening move: fewer domains, each with a clearer permitted conclusion.
Usability thresholdWhat metadata completeness and myeloid/macrophage cell-count threshold would you regard as adequate to lock the projection layer — and which resource should be dropped first if it falls short?
The gate in §1B·F is written, but its numeric threshold is deliberately left to the panel: how many myeloid and macrophage cells per patient and per condition are enough, and how complete patient/donor and treatment labels must be, before a resource is admitted.
Agreeing the threshold in advance is what keeps the inspection an honest pass/fail rather than a judgement made once the counts are known.
Novelty wordingIn light of the precedents in §2B, is the novelty wording acceptable as written — a transition-anchored, donor-aware evidence architecture, rather than a new GRN benchmark or a first macrophage network?
The literature already contains broad GRN reproducibility and benchmark work as well as macrophage/MAFB biology, so the wording has to concede those precedents without conceding the contribution. The panel's answer decides whether the proposal opens by citing the precedents and then claiming the six-part stack — donor as resampling unit, independent cohort replication, prospective lineage-control gates, curated-resource failure analysis, donor-blocked GRN stability, and bounded perturbation-context interpretation.
The specific thing to test: does this read to you as application novelty? If it does, say so here — that is the failure mode this wording is written to prevent, and it is a wording problem, not an evidence problem.
Extension gateGiven that MAFB/macrophage-identity biology is published and a TAM/MAFB regulatory-network preprint is in circulation, is the gate on the conditional extension still the right condition — and do you accept that it may never be framed as discovering MAFB or TAM plasticity?
Two separate questions in one place. First, is the §1 gate itself acceptable — a qualified direct human fate-transition perturbation resource plus a prospective causal-control rule locked before candidate inspection? Second, is the wording rule acceptable — that even if the gate opens, the extension is framed as donor-aware, evidence-qualified regulator prioritisation in this transition, never as MAFB or TAM discovery?
Also worth deciding here: whether the panel wants the professor-approved gate variant that accepts primary-human macrophage-context perturbation as constrained causal-context support, with fate-causality language still forbidden.
ScopeIs the central aim pitched at the right level — “identify and prioritize candidate transcriptional regulators”, rather than “identify master regulators”?
The framing determines what the thesis must deliver to pass. A candidate-prioritization aim is defensible on the evidence already in hand; a master-regulator aim is not, until the causal gate passes.
TitleShould the working thesis title be aligned with the claim the evidence supports, rather than with wording that implies master regulators have already been identified?
The registered project title and the permitted claim are not phrased the same way. Deciding this early avoids a mismatch between the approved title and what the evidence can support at defense.
Negative resultShould the curated TF-activity failure be presented as a headline feasibility finding, or as a methods limitation in the background?
This is the single most consequential presentation decision in the proposal. Foregrounding it makes the reliability question the thesis's novelty; burying it makes the design look like a workaround.
Causal gateWhat would you accept as a prospective causal positive-control rule, and who signs off before it is locked?
The brief requires this rule to be locked before any candidate analysis. Agreeing the rule and the sign-off in the proposal meeting is what makes it prospective rather than retrofitted.
DataWhich direct human perturbation dataset would you accept as qualifying, and on what criteria?
This is the named bottleneck. Knowing in advance which datasets the panel would accept prevents the causal stage from stalling on a disputed choice.
Stop ruleDo you accept the explicit stop rule — no candidate nomination at all if the prospective controls fail?
The panel accepting this in advance is what protects the student from being pushed toward a manufactured candidate list late in the thesis. It also defines the fallback deliverable.
ReplicationIs a 12-donor, 29,205-cell independent cohort sufficient replication for the state-axis claim?
The state-ordering result is the strongest completed piece of evidence. If the panel wants more independent replication, that changes the data-lock scope at Gate A, before formal analysis begins.
IntegrationIs paired primary-human multiome data required, or is the hybrid RNA + chromatin-anchor design acceptable?
This determines whether new data acquisition enters the thesis scope. The brief's position is that it would strengthen the work but is not a prerequisite — worth confirming rather than assuming.
GatesAre gates A–G the right decision points, and what evidence do you want to see at each one?
The gates are proposed, not agreed. Fixing the evidence expected at each gate turns them into a shared checkpoint schedule instead of a self-assessment.
Wet labWhat scope and cost are realistic for the single outsourced wet-lab validation design in Objective 4?
Objective 4 promises a specified primary-monocyte validation experiment. Its feasibility depends on resources the panel controls, and it is conditional on candidates surviving Objective 3.
Bottom line
Proposal GO. The core evidence framework is the endpoint; everything else is conditional.
The core endpoint is the core evidence framework — a donor-aware multi-omics reliability/framework thesis for human monocyte→macrophage regulatory inference. Its defensible novelty is not “first macrophage network” and not “we ran a TF-ranking tool”: it is the combination of donor-aware two-cohort replication, prospective control gates for SPI1/CEBPA/CEBPB/IRF8/MAFB, curated-resource failure analysis, donor-blocked GRN stability, bounded primary-human perturbation context, and a proposed resolving experiment.
The conditional extension to evidence-qualified candidate regulators remains available behind a gate that is currently closed, and would deliver evidence-qualified candidate regulators rather than master regulators. The pilot has demonstrated both that the data exist and why this stricter design is the honest one.
The proposed clinical-projection option in §1B is offered separately, for the panel to evaluate: a bounded downstream relevance layer, gated on resource-usability checks that have not yet been run, delivering prioritized hypotheses rather than response prediction or validated targets.
The positioning carries its own reason. Broad GRN reproducibility/benchmark work and macrophage/MAFB biology are both already published (§2B). That is exactly why the core evidence framework is the safe endpoint — it never depended on being first at either. It is also why the candidate-regulator extension stays conditional and must never be framed as MAFB or TAM discovery. And it is why “application novelty” is the wrong concession: the contribution is the stacked, transition-anchored, primary-human evidence architecture, cited alongside its precedents rather than in place of them.