Validation, governance, and leadership

GWTC-5 / O4b Lensing Incident

This is my account of a difficult moment in the LVK gravitational-wave lensing program, where I pushed for stronger statistical validation, clearer governance, and fairer treatment of the people carrying the O4b paper.

Scientific validation Unresolved pressure point
Publication strategy O4b plus O4c merger
Process concern Consultation, criteria, workload
My intervention Escalation to CBC/DAC leadership
My role O4b Lensing Editorial Team Chair
Core issue Validation of lensing statistics and backgrounds
Outcome sought Transparent standards, accountable decisions, and a healthier group culture

The GWTC-5/O4b lensing paper became more than a catalog paper. It became a test of whether the lensing effort could hold itself to a clear standard when the answer was scientifically inconvenient, politically difficult, and costly in time. I argued that high-tier lensing analyses should demonstrate their behavior on unlensed event populations before strong claims could be interpreted responsibly. Recovering simulated lensed pairs is important, but it is not the same as showing what Bayes factors, ranking statistics, or final candidates look like under the null hypothesis.

That position was not anti-Bayesian and not anti-lensing. It was pro-calibration. A Bayes factor can be powerful only when the implementation, population assumptions, background behavior, and analysis chain are understood well enough for the collaboration to defend the claim. My concern was that the group risked treating the word “Bayesian” as a substitute for the hard validation work needed to support a catalog-level result.

Why I Pushed Back

The central point was simple: a lensing candidate cannot be interpreted robustly unless we know how the analysis behaves for ordinary unlensed event pairs in the actual configuration being used.

My minimum scientific expectation was a documented validation path: calibration checks for joint parameter estimation, unlensed Bayes-factor distributions, false-alarm or background behavior, ROC-style tests, and an end-to-end Tier 1 to Tier 3 validation chain. If a full higher-tier background was computationally impossible, I argued for a smaller representative test set rather than no empirical check at all.

This became difficult because the O4b paper was already mature in many respects. Analyses had progressed, writing work had accumulated, and the collaboration needed a publication strategy. But maturity is not the same as validation. I believed that delaying or merging the paper would only be justified if the extra time came with explicit validation deliverables, not if it merely moved the problem into the future.

01 Tier 1 ranking

Broad candidate-pair screening and background behavior.

02 Tier 2 follow-up

Representative follow-up on top-ranked pairs, including null tests.

03 Tier 3 inference

Joint parameter estimation, calibration checks, and Bayes-factor interpretation.

Required bridge Lensed recovery does not replace unlensed calibration.

The Hard Part

I was not raising these concerns from the outside. I was the O4b Lensing Editorial Team Chair, and I had worked in the subgroup since before it became a formal Working Group. I had built and operated TESLA-X for subthreshold lensing searches, led the Lensing Mock Data Challenge, and contributed to the observing-run paper series. The dispute landed directly on my responsibilities, my team, and the scientific credibility of work I cared about.

Standing up for this was isolating. I had to argue against the prevailing direction of the lensing group while also trying to keep the paper moving, protect the writing team's time, and persuade division-level leadership that the issue was not a narrow technical preference. Getting support from CBC and DAC leadership required repeated explanation: the problem was not that I disliked a merger or disliked Bayesian methods; the problem was that publication strategy, background validation, and governance had become entangled.

I pushed because silence would have been easier but wrong. The collaboration needed to hear that a paper-team commitment cannot be expanded indefinitely without consultation, that validation cannot be deferred indefinitely without consequences, and that strong evidentiary language requires a defensible statistical foundation.

My position Validation first
Lensing group pressure Move the paper forward
CBC/DAC leadership Publication decision support
Editorial/Writing Team Scope, workload, continuity
The hardship was not a single disagreement. It was holding these responsibilities together while arguing that technical readiness, governance, and publication strategy could not be separated.

What Was at Stake

Scientific integrity

Candidate interpretation needed a demonstrated null distribution, calibration, or equivalent validation framework.

Publication fairness

The proposed O4b/O4c merger affected mature O4b results, short-author work, and the priority of completed analysis.

Workload

The merger changed the duration and scope of the Editorial Team and Writing Team commitments.

Governance

The decision process needed clear criteria, written success conditions, and meaningful consultation with affected teams.

The merger of O4b/GWTC-5 with O4c/GWTC-6 was not merely a scheduling change. It changed the scientific scope, leadership expectations, workload, publication timeline, and interaction with related short-author papers. I argued that those consequences had to be acknowledged explicitly, with concrete mitigation rather than informal reassurance.

Timeline

Before and during O4b

I continued pushing for end-to-end validation of lensing analyses, including unlensed-background behavior for higher-tier statistics.

May to June 2026

The possibility of combining the O4b/GWTC-5 and O4c/GWTC-6 lensing outputs emerged. I asked for transparency and argued that the decision should not preempt existing paper-team work or related short-author efforts.

Early July 2026

After a background discrepancy was reported to have a likely root cause, I continued to press for explicit success criteria and validation deliverables rather than treating a revised publication strategy as the solution.

July 12, 2026

The DAC decision was to merge the O4b and O4c lensing analyses into one paper, with a cross-division effort on background methodology and readiness expectations for pipelines.

After the merger decision

I worked to preserve team continuity, clarify commitments, and push for governance changes that would make future lensing-paper decisions more transparent and accountable.

What I Tried to Change

The deeper fight was about group culture. I wanted the Lensing Working Group to become more rigorous, more open, and more explicit about standards. Inclusivity should mean more people can contribute meaningfully to high-quality science. It should not mean lowering the evidentiary bar or leaving review expectations vague until the end.

My effort was to move the group toward a different operating model:

01

Validation before interpretation

Every major lensing statistic should come with a documented calibration or background study appropriate to the claim being made.

02

Decision criteria in writing

Publication and pipeline-readiness decisions should state who decides, what evidence is required, and how unresolved issues are tracked.

03

Respect for paper teams

When the scope of a paper changes, the workload, editorial roles, authorship expectations, and opt-in commitments should be reopened explicitly.

04

A stronger lensing group

The long-term goal was not to win a dispute. It was to help turn the group into a more transparent, reproducible, and accountable scientific environment.

Personal Significance

This incident shaped how I think about leadership. Real scientific leadership is not only producing analyses or writing papers. It is also being willing to say that the evidence is not yet strong enough, even when the room wants to move on. It is carrying the administrative burden of asking uncomfortable questions. It is asking senior leadership for support when the local process is not enough. It is protecting junior contributors and paper-team members from open-ended obligations. It is insisting that collaboration integrity matters as much as publication speed.

I regard this period as one of the hardest and most important contributions I have made to the lensing effort. I stood up for statistical validation, challenged a prevailing group direction, sought support from CBC and DAC leadership, and pushed for a more transparent standard of governance. Whether or not every point was accepted immediately, the effort helped force the group to confront questions it could not responsibly avoid.

Hardship Carrying the paper while challenging the process
Support sought Repeated escalation and explanation to CBC/DAC chairs
Change pushed Clearer validation standards and more accountable governance