Sister blog of Physicists of the Caribbean. Shorter, more focused posts specialising in astronomy and data visualisation.

Monday, 27 July 2026

Why Bother ?

It's rare that I manage to read any longer pieces on arXiv that aren't strictly about galaxy evolution, but today I indulge myself. And I'm glad I did, because this particular piece was highly provocative and well worth reading in full. This write-up will itself constitute something of a long read, so I advise getting some tea before we begin.

Ready ? Good. Here goes then.

David Hogg's self-proclaimed "very white" paper thankfully isn't so in the horribly racist sense, but in the "here are some semi-organised thoughts that might be worth considering" sense. It's entitled "Why do we study astrophysics ?" and it's ostensibly written in the context of ever-increasing AI development. His goal is to provide a stepping stone to understanding when and how LLMs should be employed for astrophysical studies. He doesn't attempt to come up with a final answer to that question, because conscious or not, the effect of a machine that can answer questions more accurately than human experts still represents a technological and sociological singularity. Seeing beyond that point is too big of an ask. 

Instead, he tries to tackle the fundamentals of why we do our job at all, thinking that this is a necessary precondition for how and why we might go about automating parts of it. If we can figure out why we're doing what we do, maybe we can better understand what we should do given expected technological progress.

As an essay, I found this one thoroughly excellent. There's much here I disagree with and much I support, all of it well-argued and clearly stated. It sticks to its central theme but covers a very wide array of topics along the way. There's actually not too much about AI in here, the focus being more on the human side, but I think it may have at least a germ of an answer as to how we'll proceed come the killer robot uprising technological singularity.

To be fair, Hogg's essay isn't the most linearly organised piece in the world, often feeling like something of a memoir. I've tried to keep this summary-cum-commentary to the linked themes and bits I thought I had something worth contributing to; I've deliberately avoided issues where I disagree but don't think the argument would get us anywhere (such as whether LLMs are truly thinking or not... this is very interesting to me, but makes no difference to the arguments here). 

First, I'll look at the main thrust of the article : what astrophysics is and why we do it, which involves various thoughts on how we go about this. Then I'll conclude with a much shorter section on what this might mean in that future where we have vastly more powerful analysis capabilities than anything we possess today, but which we can reasonably expect to have in the coming years.


1) Astronomy Today

The professionalism of science 

Much of astronomy is now done using large facilities by enormous groups, and the way these operate is inevitably different to small groups in a lab in someone's basement. Here I think the word he's searching for is industrialisation, not professionalism. The career-based nature of astronomy has already been long established, but the escalation of scale is still relatively new. The point is that you can't have people just casually mucking around on dedicated survey tools or slapping their own instruments onto billion-dollar facilities. This was absolutely possible in the old days – the underside of the Arecibo dish was littered with discarded receivers – but this kind of approach is all but dead already.

Astronomical data production is becoming extremely professionalized, and in a very particular way. Astronomers, in the case of Gaia, are just end users; end users of curated, calibrated data, delivered by a combination of the ([military-built] secret) spacecraft and the (absolutely great, professional, and open) DPAC... it wasn’t built or operated by astronomers.

Even university-based projects endeavour to produce science-ready data products that can be queried through application programming interfaces, plotted, and analysed without much worry about where they came from or how they got here. I have been involved in bringing about this change, and in many ways it is absolutely great. It democratizes astronomy, since it lowers barriers to entry. It creates an open-science space, in which every project benefits from the output of every other project.

But it does have a strange consequence, which is also related to professionalization : for some kinds of projects in astrophysics, there isn’t a huge difference in capability between a classically-trained astronomer and a newly trained data scientist... a data scientist who has taken an astronomy class might be better prepared than an astronomer who has taken a data science class.

Which is relevant, of course, because LLMs can absolutely do data science.

I think this kind of development is only partially inevitable though. We need big data and big data needs big facilities. But we also need small data, that is, data we can analyse in extreme detail. That may still need big instruments but it also needs small groups, and small groups can, and should, continue to operate in their current fashion. Analysis of big statistics is surely going to change, but on smaller scales, perhaps, is a realm where we can still do the low-level stuff ourselves. Here is where we can invest considerable time doing some aspect of the data reduction and analysis by hand and still have a good chance of making useful, interesting discoveries.

My reasoning is purely pragmatic. Manual data reduction and analysis can be extremely time consuming, but is perfectly manageable on small data sets. We've already abandoned this approach for the largest data sets because it's just not possible to do them by hand. But I fervently believe there is a great deal of value in learning to do the low-level stuff (the hard bit), even if later you never do it again. As a general rule, the better you can operate without a specialist tool, the better you'll be when you get to use it. Maybe for humans, then, the future lies not in big data, but in small data, in extreme specialisation rather than generalisation.


What is astrophysics ?

It's that which produces novel information about the Universe, says Hogg. Reading about it doesn't count, you have to do actual research. But... you also have to document the results. Unrecorded data doesn't contribute to the pool of knowledge from which others draw, so if you don't document it, you might as well not bother. 

Astrophysics, then, is the literature, in his view, and it's this act of producing literature after a novel investigation which best describes the process of doing astrophysics.

I think it's hard to dispute this, but he has a couple of other points which might be more controversial. One is that software isn't as important as the results it produces. I actually do agree with this, because he explicitly declares that software should have associated papers. This then makes it just as important for certain metrics as actual science, and I think it's absolutely right that the effort of software development be properly recognised. 

What he means here, I think, is only that software itself is not science. We write software so that we can analyse data, not because it's intrinsically worth having. Software which isn't used is as pointless as data that's not recorded.

He also notes :

Astrophysics, like any science, contains a lot of “implicit knowledge” or folklore about things like how to observe, how to reduce data, how to organize projects, how to visualize data and models, how to read and write, and so on. Much of this never appears in the literature. Is that not also astrophysics ? Yes it is, but it is astrophysics practice. The results of astrophysics — the scientific conclusions and debates — are in the literature.

Yes, but my answer here would be that we should absolutely record as much of this "folklore" as we possibly can. Some sort of journal of astrophysical methods – not describing mathematical procedures, but the really low-level stuff of what to do with the data and how to interpret it – would be valuable, I think. Such papers wouldn't have the lasting value of results papers, but they would make a lot of people's lives a lot easier. Implicit knowledge should be made explicit wherever possible*.

* Though it is ultimately impossible to record literally everything. Some things you simply have to do.

His second more controversial comment, with which I vehemently but provisionally disagree, concerns papers as a metric :

The second comment is that I often hear software (and hardware and engineering-oriented) people say that they “have to” write papers because papers — and the citations that they generate — are “the coin of the realm.” Papers (and the authorships on those papers) and the citations of those papers are not “coin” of anything ! They represent our recording of what happened, what we learned, what we know, and how we know it. Citations deliver provenance, not reward.

I'd love to agree with this but I can't. In terms of career advancement, it's not sensible at all. Like it or not, astronomy as a professional/industrial career does have certain requirements common to all employment. We need to get paid and we need to ensure job security, and we can't and shouldn't ignore this. The era of the gentleman-scholar is long over : I mean, sure, I'd love to give everyone tenure and a sack of money, but until we do that, papers absolutely and undeniably are the coin of the realm.

In fact, even in terms of strict science, I still don't think I can agree. If astrophysics is the literature, as Hogg claims, and our goal is doing science... surely it's literally true by definition that papers are the "coin of the realm" : inasmuch that if we should be judged by what we produce at all, this should be the primary means by which we do so. I can't get my head around the alternative, which would be a bit like saying that we shouldn't judge a painter on either the quality or quantity of the paintings they produce.


People are the ends, not merely the means

What might explain the above difficulty is what to me feels like Hogg's most controversial and complex point, one which I'm allocating three subsections to examining.

People, says Hogg, are what astrophysics is really all about. It's not about the Universe at all (and he's emphatic and explicit on this point, on which more below), it's about enriching ourselves.

When we employ a graduate student to perform some work, it absolutely must be because the graduate student will benefit from that work, not merely because that work needs to get done. I have heard it said, more than once, in research contexts, that an LLM can do some task “better than a graduate student.” That language makes me uncomfortable, because it is taking an extremely instrumental view of graduate students. Are graduate students in our groups and our laboratories and our universities to do work ? Or are they here to learn ? 

We train PhD students not merely to amplify our own research programs, but to create opportunities, and specifically opportunities for them. Every person is a human being, whose personal development is more important than our short-term scientific accomplishments.

I mean... sure, to a point. I think it's the old Platonic point about whether education is about discovery or change, and it can only ever be both. We shouldn't be using graduate students as literal tools to solve the problems we want solving; they are not there to do the boring grunt work that needs doing but we don't want to do ourselves. But at the same time, we should be getting work done. We should be solving problems ! We should be learning about what Nature is, not just endlessly pontificating on what it might be. 

The only real solution here, I think, is to find graduate students who share our interests so that the result is mutual benefit. We should be giving them problems that both advance knowledge and advance their own abilities. Otherwise, we risk running into one of Plato's weirder quotes (Republic, book VII, 503b) :

Then if, by really taking part in astronomy, we’re to make the naturally intelligent part of the soul useful instead of useless, let’s study astronomy by means of problems, as we do geometry, and leave the things in the sky alone.

I do think it's important to realise that the people doing the work are first and foremost people, not tools. The further we can get away from the mentality that productivity is the only result that matters, the better; the more we can suppress the need to be competitive, the more we can suppress the idea that we need to live to work – even when that work is something we really enjoy – the better the situation will be for everyone. 

But two important tangential points crop up here. First, Hogg declares that this means that not citing relevant papers is (ignorance aside) actually unethical. 

It is ethically required that our papers cite the work that is relevant to the work we are doing. You can’t decide not to cite a relevant paper because you don’t like the author, or don’t like the author’s institution, or don’t like their funding sources. In particular, if the literature gets flooded with work of relevance to your research program, you have reading to do, and citing to do.

I object very loudly to this ! And not just because I'd read dozens of papers that damn well should have cited me but didn't. First, pragmatically, the increasingly industrial scale of astronomy literature means that reading every paper on a topic is not a sane choice. As per the last post, reading hundreds of pages of largely-irrelevant text (and astronomical papers tend to be extremely dry, which is not a minor point) is going to result in negative value, not merely slowing things down. So no, exactly for the sake of not treating people like instruments, you absolutely do not have to read and cite absolutely everything. That's mentally destructive, not productive. You don't have a duty of self-destruction.

Secondly, from a moral viewpoint I also disagree. I see nothing at all wrong in deliberately not citing papers where we don't find the results and/or methods credible, and might even object if I was compelled to cite something I didn't believe – and I'd certainly have a very big problem if I was told by a reviewer to make something sound more plausible than I thought was really the case*. This is not to say we should avoid controversy and it certainly doesn't mean only citing the things we agree with. It only means that we don't have to cite the whole history of a research program and give equal weight to every long-discredited idea or failed avenue of inquiry.

* Hogg actually says himself that you can't cite work you don't trust, but seems to think it's obvious that we can trust people and can't trust LLMs.

There's one aspect here with which I do, however, violently agree :

Every scientific paper is written to help all of its writers, and all of its readers, learn and grow, no matter their career stages.

My take here is not about the content of the paper so much as their style. We need to write for each other, as human beings (which Hogg does very well indeed), not as automatons who require total clarity and unambiguity. If you want me to cite more papers, reform the standard requirements for a manuscript. Make them shorter, better organised (results first, then detailed methods) and more readable (allow the occasional joke, stop being anal about contractions and punctuation FFS). But this is a well-worn hobby horse of mine so I'd best not go down that route again today. 

Hogg goes further and suggests that grant funding to hire students to do work might also be unethical ! And again, I cannot agree with that. We're not running a charity and it's not at all wrong to expect productive output (though we absolutely do need to be flexible in our expectations of that output). This is also in stark contrast to his later claim that we need to use our resources efficiently and get correct, rigorous results. We should seek good working conditions, but ultimately we are doing work. The results do matter.


The answers don't matter

Here we come to the heart of the problem. Hogg genuinely believes that the results of our research aren't important. In one of his oddest moments, he says that if we really cared about the results, we wouldn't do astrophysics ourselves but pay other people to do it for us... this is weird every way I look at it. I just don't think that's how people work, because extending that reasoning, nobody would ever do anything for themselves at all. And of course, it's a pretty perfect example of so-called effective altruism, which Hogg calls an "absurdity" ! He's not wrong about that, but my goodness me, the contradiction is as a glaring as glaring can be.

This baffling oddity aside, Hogg's argument is not to say he thinks we should all quit and do something else. His claim that the results don't matter is more specific and more strict than that... he think the investigations are worth doing, that that's where the benefit lies – in improving ourselves – but that what's actually going on in the Universe is of no consequence to what's going on down here. 

That's, err, quite the hot take there. But it deserves more examination.

Hogg has this highly annoying phrase, "clinical value" which he best expresses thus :

I like to say that the sciences have a “left edge” which is about fundamental understanding, and understanding for understanding’s sake. They also mostly have a “right edge” which is about what I like to call “clinical value” but you could call application or use in the world for technologies or policies. 

I claim (and maybe this is a bit controversial) that astronomy has no right edge. That is, there are no useful things in the world that flow from astronomical discoveries and results. I have spent years of my life estimating the comoving volume of the Universe, measuring the local dark-matter density, and finding planets around other stars. No human outcome or pragmatic capability has been affected in the slightest by any of my results. Literally nothing helpful to humanity arises here.

No sir ! No, I won't have it. First, the reasons why we do astrophysics – the whole title of the essay – are to me obvious. I cannot understand people who don't have any interest in understanding the nature of the world in which we live, and for those that do, then understanding the most miniscule corner of it and ignoring all the rest seems like a clear sign of insanity. For me, observational astrophysics is absolutely a fundamental science, more so, I would argue, than theoretical physics : that's just making up a bunch of stuff, which is valuable, but ultimately tells us nothing about reality, at least not with any certainty. It is observation alone which can do that.

Knowledge of what's beyond the sky is not some abstract wishy-washy thing, but essential in understanding the truth of our own existence. Would it not matter if the stars were holes in the curtain or night rather than fusing spheres of hydrogen ? Would it not matter if the nature of reality were that we were actually inside a giant koala rather than an immense vacuum ? I think it would, and in fact it might well form the basis of all our other knowledge.


Astrophysics is useless

Which leads to the final part of this section, and the second, closely-related reason I think we do astrophysics. Hogg claims that it's not for spin-offs and these don't count as the "clinical value" or right edge. With very few possible exceptions such as discovering dangerous asteroids, and in previous eras understanding chronology and navigation, he claims that these aren't the reasons at all :

Nothing in the world of things or people hangs on the precise value of the age of the Universe. Astrophysics may occasionally and accidentally produce something useful. But astrophysics is not done with the goal of obtaining clinical or practical value. A science has a right edge if and only if the associated clinical work actually tests or exercises the specific results of the science. None of astrophysics is justified in these right-edge terms. No astronomer (that I know) is improving the calibration of JWST instruments because they want the US Navy to have a higher kill rate.

No ! Astronomy's right edge is not in "the clinical value [which] lies in its feeding of humanity’s love" or some other airy-fairy thing that Hogg justly raises as failed counter-arguments. It lies in telling us what is not true. It defends us against ignorance, and the price of ignorance can be extraordinarily high. It becomes extremely difficult to maintain that you need to sacrifice people to appease the gods of the sky when you realise that there simply aren't any. Cosmology has direct moral implications : just because we no longer take a direct moralistic approach to cosmology, as was done throughout medieval history and earlier, and as Tolkien did brilliantly in fiction, it doesn't mean that our morality isn't affected by our understanding of cosmology.

Now to be fair, the precise values of different parameters do not always constitute such a hard right edge. It's not obvious how the exact distance of Proxima Centauri or the HI content of the M31 galaxy could have any moral value whatever. But collectively, we need all these incremental findings to get to the good stuff. We need the flies in the ointment to break our understanding and shatter our conceptual frameworks every once in a while. We need things like the perihelion of Mercury to tell us that Newton is wrong and time and space are themselves not at all what we thought they were, the full moral implications of simultaneity breaking still being something we haven't got a handle on. And we only get those results through slow, careful, methodical measurements.

Does it matter that those of us working on astronomy aren't doing so for the hope of such a breakthrough moment ? Does our motivation being purely intellectual simulation invalidate this hard right edge I've suggested ?

No, I don't think so – not at all. It's true that many of us like the pure research side of things, that we do our jobs (in part) precisely because we can avoid having to be responsible for other people. I too like the fact that nobody's daily lives are at all likely to be impacted by the velocity width of a dwarf galaxy I publish deep in a table of a paper that hardly anyone will ever read. But this does not mean the work isn't worth doing for its own sake, that it won't potentially contribute, albeit in a small way, to the revolutions in thinking which will eventually and inevitably follow. And those who are working with more express goals – if there actually is anyone out there calculating the distance to Proxima Centauri purely to refute astrologers or the hope of fortune and glory – well, more power to them, and equally, their motivations aren't invalidated by the pure interest sake that the rest of us pursue.


2) Astronomy Tomorrow

How does all this mean we should prepare for an astronomical future in the age of AI ?

Hogg proposes two extremes, both of which he views as undesirable. One is that we hand over everything to the LLMs and literally let them do everything, or at most, we try and curate their findings to sort the good from the bad. This would seem to be a pointless exercise in which we don't ever really learn anything, we abandon the joy of the process and reduce ourselves to mere instruments. Pretty much nobody wants that. 

See, I think Hogg does have a point that the human element matters : we do astrophysics for our own enrichment and reward, and both the process and the results matter to us. Even if an LLM had such emotional motivations, there would seem to be self-evidently no value whatever in letting them do everything. That'd be like sending someone to go on a rollercoaster on your behalf. Having the experience, not just having casual access to the results, matters.

When we offload that work to LLMs, we are no longer doing astrophysics, we are no longer becoming astrophysicists, and, eventually, we no longer are astrophysicists. The let-them-cook policy, in the end, leads to the death of astrophysics, the end of astrophysics at universities, and the end of astrophysics education. Astrophysics would no longer be by humans, and then it would no longer be for humans.

The second extreme is that we ban LLMs altogether. This Hogg views as bad because LLMs can be genuinely useful, the effort to seek-and-destroy LLM content would be hugely inefficient and wasteful (again, we'd become mere instruments), and telling people how they can and can't do their research self-evidently violates their freedoms.

Hogg's tentative and intriguing suggestion for a middle route is that we treat LLMs as colleagues who aren't part of our own team :

You might ask your colleague for help finding something in the literature, but you wouldn’t ask your non-coauthor colleague to write the introduction of the paper you are writing. You might ask your colleague to help speed up your code, but you wouldn’t ask your colleague to write your code. 

I quite like this. Necessarily, the middle route must be allowing the LLMs do some of the work rather than all or nothing, and the essence of having some simple guidelines for good practice is sensible. 

I don't think it will work out exactly like this though : recently, I've been "vibe coding" a quite elaborate program and I'm convinced this is indeed the way of the future. I've resisted this practise for a long time, but decided that there was one particular code I really wanted to exist that I didn't have the time to work on (to be shared in a future post). Vibe coding is not zero effort, far from it, but it does work. I very much doubt that we're going to insist on coding by hand for too much longer, any more than we insist on adding and dividing using pencil and paper. Still, the basics of Hogg's idea are interesting.

I need to finish with a few assorted caveats :

Another idea is that, given our respect for our readers, we shouldn’t ask them to engage with something that took way less time to write than to read.

I don't think so. The content isn't any the less valuable based on effort. What seems obvious to one person can be profound to another... I mean, I heard a podcast of Mary Beard – Mary BEARD, for crying out loud – dismissing Marcus Aurelius' Meditations as of "no value". So much for that.

A second is that of the nature of LLMs : according to Hogg, we cannot trust their output, the text isn't meaningful until a human reads it, LLMs aren't reproducible, they currently only produce slop, and they can't take responsibility. All of these I think are only partial truths : we can apply the same methods of trust as we do for humans; the fact that humans will always have to read the text would seem to make the argument that LLMs lack any true understanding to be largely pointless; I agree that LLMs are not deterministic but this isn't quite the same as lacking reproducibility; the idea they only make slop is simply a garbage claim; and much more development is needed here as to what we actually mean by "taking responsibility".

 And finally one throwaway, tangential claim I cannot responsibly let slide : 

Of course it is important to remember that the human practice of astrophysics, at least in its current form, is also very damaging to the environment.

Without reading the citations provided I instinctively think this statement is of negative value. There's no way that astrophysics represents any sort of significant environmental problem. The kind of practises which are truly damaging are those of big businesses, industries, and the exploits of billionaires. It's right and proper that we try and set an example and constrain our environmental footprint. But we should do so only insofar as this helps curtail the much worse damage that's being done by others. Otherwise we risk falling into a Calvinist sort of pointless guilt, in which all we accomplish is to feel horribly depressed for expending energy on a Zoom call while billionaires continue to fly first class across the Atlantic for weekly holidays. 


Conclusions

Phew ! Well done if you made it this far.

I wanted to try and venture a few thoughts as to what might happen next. But, having rewritten this section several times and always ending up with content I never quite believed, I decided to abandon this approach. Instead, I can maybe offer some thoughts about why these predictions are so difficult – and maybe just a hint of something more.

The obvious reasons are that trying to predict what something more capable than ourselves might achieve is fundamentally difficult, and of course the rapid pace of development makes prediction inherently uncertain. A slightly more interesting factor is that different people will adopt different approaches. Some will despise AI, some will love it, some will see it as a tool, others will use it more like collaborators. That we already see this happening makes giving any one answer about what's going to happen next a flawed question, like insisting that there can be only one answer to "what happened in Britain after the Romans left ?" when in fact there are many.

Related to this, I also think that people have very different ideas as to which part of the analysis they'd like to automate away. For me, visual inspection of the data is the fun part. For others, that's the bit they most want to avoid and they want to concentrate on the mathematics or the hypothesis-generation. So prediction difficulties are hit by a double-whammy (at least !) on inhomogeneities : people have different attitudes to AI and different preferences to what they want to automate. Maybe it'll all just balance out.

Another difficulty is feedback : that we don't know how this level of automation will affect us. I'd like to think that it'll be linear, that we all pull back on the stuff we don't like (different though that will be for everyone) and concentrate our mental resources on the stuff we do. That is, our mental capacities won't diminish, we'll just redirect ourselves. I think that's probably likely to be the case most of the time, because while Hogg takes it too far, people in academia do value the experience of problem-solving in itself. They aren't likely to want to just stop doing that – indeed, some of them might even not be able to. But we also have to consider the temptations towards laziness, to jump straight to the final answer... and more insidiously, that maybe only the low-level stuff is sufficient to really keep the brain working at peak ability or prevent it from degrading. 

People on different sides of the AI debate all have very different views on this. I lean towards "this will just be a good thing, most of the time". My personal experience is this is something which really lets us get shit done, and nothing remotely comparable in getting-shit-done abilities has preceded this in my lifetime. Not even close. It's hard not to be optimistic about that, and I incline towards the view that a thing which is good for getting shit done is highly unlikely to actually decrease the amount of getting shit done. Even so, I don't think the effects are fully predictable, and I don't dismiss the tendency to skip the legwork* and get to the answer instead.

* Isn't skipping itself legwork though ?

I also have to recognise the special privilege of astronomy here. It's not quite that our results are of no importance, as Hogg claims. It's a question of precision and timescales. On the long term, our broad results are as important as anything else in any field of knowledge. But the exact values we determine in the short term are indeed of no importance to anyone else except ourselves. This frees us from any immediate need to be productive : our findings almost certainly won't cure cancer or relieve pain or solve famine, except possibly through spin-offs. So this means we can, and perhaps will, continue to do some of the low-level stuff genuinely for enjoyment, just as I've written this extremely long blog entirely by hand* because it's something I wanted to do, not because I think more than half-a-dozen people are likely to read it or even because I thought I could do a better job than a chatbot.

* I deliberately added a keyboard shortcut to make typing en dashes easier, so don't let those fool you. I happen to like en dashes, mmkay ?

The importance of small data in astronomy remains key to allowing humans to make genuine contributions, just as amateur astronomers can and do still make valuable discoveries for the professionals. This is not something that has direct equivalences in other fields, but it does give us some clues to the (short-term) future. While Hogg may be right that an LLM can write a paper 100,000 times faster than a human, they don't have any innate desire to do so. There's no topic an LLM actually has an interest in because they literally don't exist until prompted by a human : they have a crude agency, but no consciousness. And yes, while it might eventually be possible the deploy the large-scale compute Hogg hypothesises* could lead to factors more in the billions, where we could simply ask, "Please solve all problems in astronomy" and get something back that wasn't drivel, I will go so far as to say this isn't happening this decade.

* If I were him, I would make it my professional mission to have a hypothesis named after me.

So we're safe for the foreseeable future. The techbro predictions are hype, but they're not made up of nothing. AI is and will be transformative for astronomy. Anyone thinking that we can ignore it, that things will carry on as normal, or even that they can clearly see where this is going, well, enjoy your blissful ignorance, but I'm afraid you didn't get the memo. My only prediction is that the solution will be obvious after the fact and a handful of people who got lucky with the right call will proclaim themselves wise sages... unless they too are replaced with killer robots. Only time will tell.

Tuesday, 16 June 2026

AI Can Help Us Publish Less, Says Scientist

Not really all that much about AI in this one, actually.

I think everyone agrees that "publish or perish" is bad, but I don't think the approach suggested here makes much sense. However, I do like the following point very much :

AI is entering a publication system already swollen by proliferation, marked by signs of declining disruptiveness, and under growing pressure at the level of review and evaluation. Under those conditions, scientific papers can acquire negative epistemic value: not because they are wrong, but because the understanding they add no longer compensates for the time and attention they draw away from editors, referees, readers and colleagues trying to place them within what is already known.

Once papers and citations become central to hiring, promotion and funding, while publication itself also becomes a commercial object,  proliferation acquires a force of its own. Science fills with large bubbles of urgent-but-not-important writing: work that is timely, legible to evaluators, easy to package and profitable to circulate. 

When scarce time and judgment are drained by papers whose contribution no longer justifies what they demand from everyone else, they contribute negatively to the collective production of knowledge, slowing and hampering the rise of the mountain.

Well, I agree ! The author is also careful to say that having lots of incremental papers is not in itself a bad thing. But we're reaching the point where trying to maintain the vast wealth of relevant knowledge needed for a small amount of progress is outlandish. We require typically 15-20 pages or more to describe in meticulous detail what was done, why it was done etc. etc. etc. all for the sake of a small advancement that nobody will care about except a handful or direct competitors who will tear it apart limb from limb... we're burning the candle at both ends, making an enormous amount of work for ourselves both when reading and publishing.

I exaggerate, but only slightly. We also have to spend an inordinate amount of time adhering to strict and and utterly pointless journal standards which do exactly nothing to advance the state of the field and do an awful lot towards making the final paper less readable.

So I agree with the diagnosis. I'm much more skeptical of the suggested treatment, not because I have anything against AI in science (quite the opposite !), but because I don't think this is right approach and won't do anything much to address the problem.

The first concerns the visibility of non-paper contributions, such as code on GitHub, data on Zenodo, curated benchmarks, public notebooks, reproducibility packages, and living syntheses. They have existed for years and are valued by practising scientists. But they have remained second-class citizens: largely invisible to hiring and evaluation committees, and usually legitimized only via a paper that describes them rather than recognized in their own right. AI can change that communication layer. 

Can it though ? Only very weakly, I think. It can make a code or other product more intelligible to evaluators, but this won't mean anything if they don't have some box to tick on their reports.

The second concerns time. AI can absorb much of the routine labour that now consumes researchers’ effort: literature mapping, code scaffolding, documentation, reproducibility checks, exploratory work on alternative paths, first-pass synthesis across neighbouring literatures... It is that it can remove some of the weight that now pushes many worthwhile directions out of reach.

I agree with that one. AI – and I mean AI from the last few months or so with very low hallucination rates – is very, very good at turning routine but unintelligible research into something accessible and useable by experts who aren't specialists in that particular field. It can also be trusted at least with grunt-work that produces easily-testable results. Don't want to spend time writing a GUI for your enormously complex script ? I completely get that, it's boring. But you can easily have an AI slap something on which, if not perfect, is still massively better at not having one at all, and can generally be refined to something of a decent standard fairly easily. Thus you end up producing stuff which is not only powerful, but actually useable and accessible to a wider audience instead of the hardcore loonies who insist that everything should be done via the command line for some reason.

The third concerns evaluation. AI can strengthen the front end of review by helping editors, evaluation panels, and funders with triage, novelty checks, literature comparison, technical consistency checks and the detection of obvious pitfalls... This is one of the places where AI can directly counter the danger of negative epistemic value. When scarce time is spent processing papers that add too little in return, knowledge suffers. Review would still take time where needed. What could change is the amount of low-level labour surrounding it, so that more of the community’s limited attention is reserved for contributions that genuinely deserve it.

Here I think this is basically true, but not so much for review itself as for distilling knowledge into what researcher's actually need. I still want a human expert reviewing the paper and checking the whole thing carefully for errors, but once published, AI is very powerful for checking on whether a paper actually contains something I actually need to use. It's just not reasonable to expect authors to fully read hundreds of 20+ page papers in full when producing their own; the vast majority of citations are selected only because of one or two key results in each paper, not because everyone is reading absolutely everything.


I've said it before and I'll say it again. What we need are two main changes, one at the level of the journals, and the second at the level of evaluation. The two are inextricably linked. Instead of publishing just in a regular journal or Science/Nature (i.e. ordinary versus prestigious), we need far more journals and divisions within journals. We need to actively demark papers that required a shittonne of work from those which were rattled off in an afternoon. There is real value in disseminating pure ideas with absolutely no testing, but such a paper shouldn't be held to the same standard as the results of running a huge simulation or cataloguing an enormous set of observations. And we need, therefore, to insist that these different levels of papers – which need different standards of review rigour, clearly and publicly stated (most journals do nothing of the sort, never stating what the reviewer is actually supposed to do and when they should shut up) – are actually accounted for in evaluations. Maybe your department already has lots of hard-working incrementalists and needs someone more creative. Maybe it's the opposite. All have value in the right context

(Incidentally, my institute does account for non paper-producing duties in our internal evaluations, but I've yet to see this much used in external applications like jobs and grants)

A closely related point is that we probably need to think more about how we want papers to be structured. The prestigious journals tend toward a much more readable format : here are the key results together with the primary reasoning and potential pitfalls in this 6-page report, and here, in this 20+ page appendix, are the deep technical details of how we did it. This makes it massively easier to read the key results if you don't need the gory details, and has no real downsides if the technical stuff is what you're after.

So : papers which are easier to read; papers reviewed to different standards and with different labelled metrics; and evaluations which account for the different value that different types of product bring to the table. That would help a great deal, I think : not so much by publishing less as publishing differently and recasting what it is we actually have to read. AI might help here, but only as a second-order effect. It's not the main route, in my view, to beating "publish or perish" culture to its deserved death with a big stick.

Wednesday, 18 March 2026

The Secret History of Dark Matter

One of the nice things about doing all my academic reading on a digital tablet is that I can download papers I'd like to read purely for the sake of interest. I don't get much chance to actually read them, but it's better than having a huge list of bookmarks I'll never check, or a stack of printed papers so large it could qualify as a carbon sequestering facility.

Finally, I managed to get round to reading one such paper, and it turned out to be a thoroughly worthwhile read. Maybe a bit on the lengthy side, but then that's what this blog is for.

Anyway, the popular narrative history of dark matter goes something like this. Jan Oort and Fritz Zwicky made some early claims that they might have found it back in the 1930s, but it was all from purely weak observational evidence. There wasn't any particular theoretical reason for it, so everyone ignored it until the 1970s when Vera Rubin and others started finding that galaxies were rotating much too quickly. Et voila, paradigm shift, everyone got very excited, and this resulted in the modern cosmology we know and love.

This paper* makes some important revisions to what actually happened. Rather than being pure observational luck, Zwicky had clear theoretical motivation for dark matter. Some of his underlying reasons for expecting dark matter have long since been thoroughly refuted, but some have intriguing similar aspects to modern theories. And he was even, quite likely, actively searching for it, with his technique being one that's still in regular use today. His sample size was absolutely shite – seven or eight galaxies would never be enough to convince anyone, but the method was sound. 

* I'm unsure of the provenance of the article. It appears to be only uploaded to preprint servers with no hint of whether it's submitted to a journal or not. 

What he did not do was stumble on a result he couldn't explain and invent the idea of dark matter as an ad hoc "fudge factor". More on this later, but even from its earliest days, there were already strong theoretical reasons to believe dark matter existed before observations started to get ahead of the game. Today, everyone knows about Zwicky, but most people forget the theoretical paradigms in which he operated. This paper attempts to set the record straight.

What follows is my summary of the paper. I've tried only to shorten and simplify the content rather than put too much of my own spin on things.




Zwicky did not at all like the idea of an expanding Universe. Today, this is as well-established as any result in science can be, but at the time, he had good reasons to be skeptical. The difference in redshifts caused by expansion of space and motion through space were not yet fully appreciated, and the "breathtaking speeds" of galaxies moving at thousands of kilometres per second therefore seemed ludicrous. From a contemporary vantage point this seems weird, but when you only have a meagre handful of data points and they seem to be implying that something outlandish is happening, most of the time it's actually quite sensible to bet on your pre-existing ideas. 

This is what led to Zwicky's "tired light" hypothesis as an alternative to cosmological expansion. The idea was that photons would lose energy as they travelled for long distances, becoming redder and redder. Galaxies might indeed be at stupendous distances – I see no indication Zwicky ever doubted this monumental discovery from Hubble – but their speeds might be an illusion. 

The thing is, what would cause a photon to become tired ? Zwicky's answer was that there must be some intervening material, unseen through direct observation, but inferable through its effects on photons. And he wasn't the first to suggest that there could be some quantity of dark matter out there, with the authors suggesting that actually the opposite hypothesis – the idea that all matter must be luminous – was regarded as equally audacious.

The prehistory of dark matter is long and complicated. And it really is prehistory, because it really seems to be only much later that we get to the idea of a genuinely new type of substance, the concept of a material that only interacts with our own through gravity. None of the earliest ideas ever suggested it was anything other than matter which was perfectly normal, just in a state where it was bloody difficult to see. This includes things like Mitchell's "dark stars", which generated photons but which were trapped by the star's massive gravitational field; unseen planets like Neptune that just hadn't been spotted yet; most crucially of all, Einstein and de Sitter's cosmology required dark matter to maintain a flat Universe.

So Zwicky's motivation for dark matter seems to be neither quite that he was the curmudgeonly contrarian of popular lore (though he definitely was a cantankerous git) nor a visionary ahead of his time. Rather he was operating in conditions which are not really directly comparable to the modern scientific view at all. He had three main motivations for expecting dark matter, and these are best understood on their own terms. First, he thought this could explain away the expanding universe, which he viewed as a problem rather than a reality*, through his tired light hypothesis. Second, he knew dark matter could greatly help with the Einstein de Sitter (EdS) model, which Zwicky seems to have favoured. And third, he also thought it could explain the origin of cosmic rays.

* This is much my own view on dark matter. I don't quite understand why many people treat it as apparently obviously problematic and in need for explanation, rather than accepting that this is just what the data shows. I'm simplifying here, but you get the point.

Tired light is well known, and long since refuted, but the problem of cosmic rays is more often forgotten. Nobody understood where these high energy particles were coming from, but since they appeared to be uniform across the sky, they either had to be from something very nearby or very far away indeed. Since there were no obvious nearby sources (that is, within our own small patch of the Milky Way) that might explain them, a cosmic origin was very reasonable. Not knowing about active galactic nuclei and the like, Zwicky's dark matter seemed like a pretty good site for their origin. I have to say I don't quite understand the author's argument as to why Zwicky didn't believe they could originate from anything luminous, but still, if they originated from low density matter filling all of space, this would naturally explain their uniform distribution across the sky.

He also had a good reason to search in clusters. Lord Kelvin had suggested in 1904 that galaxy dynamics could give clues to their total mass independently of their brightness, and Zwicky extended this to clusters (incidentally he also realised that, in partial contradiction to Einstein, a galaxy cluster might be so massive that gravitational lensing there might be so strong as to be detectable).

This meant that Zwicky had very solid grounds for targeting Coma. He doesn't set out his motivations in his own paper, so all this is a contextual reading by the modern authors, but in my view it's a compelling one. Zwicky had good reasons to believe dark matter could be detected in clusters by examining their dynamics, and similarly convincing arguments for its existence which were completely independent of dynamical considerations. It's entirely credible that his observations were done deliberately for this very purpose, and his putative discovery of dark matter wasn't something that he just happened to have stumbled on by chance at all. 

The most interesting aspect of this to me is that dark matter was a prediction of relativity. I've long wondered if an early detection of dark matter would have throttled relativity in the cradle, but the answer from this paper is a clear "no"... but not for reasons we could still justify today. The thing about relativity is that it required dark matter for a flat Universe, but the amount required would have been far larger than in our modern estimates. Today's cosmology uses dark matter to explain the dynamics of galaxies and clusters, but the cosmological constant (and inflation) to keep the universe flat. At the time, it seems the equivalent mass density of the constant wasn't considered to be sufficient to do the job of flattening space. 

And it's important to remember that Zwicky was still making one hell of an extrapolation. From his seven or eight galaxies, he assumed that the rest of the Coma cluster galaxies (hundreds strong) were in stable equilibrium, and thus derive a value for the total mass which happened to be in agreement with the predicted dark matter content needed for the EdS model. But given that even Hubble's constant was not at all well-constrained at the time (Hubble's first value as 500 km/s/Mpc; today we think it's around 71 km/s/Mpc), the error bars on this were massive. So again, relativity predicted dark matter, but not directly, and not for reasons we can now sustain. History turns out to have been more complicated than my counterfactual musings.




The final part of the paper is more philosophical, considering whether Zwicky's idea really constitutes an ad hoc hypothesis. Certainly it seems not to have been something he just invented on the fly; he may have even been deliberately searching for it. As far as Zwicky's particular idea goes, the answer here is a decisive "no". He may very well have been doing the classic scientific model of hypothesis testing, with his observations set in a clear theoretical framework – and was definitely not trying to save Newtonian gravity from relativity, as has been claimed.

What about ad hoc hypotheses more general ? The paper gives quite a thorough discussion on the different perspectives on these, noting that the existence of Neptune was arguably just such a case. Explaining the orbits of the other planets was difficult without an extra one hitherto undetected, but there was no other good reason to expect the existence of such an object. And of course Neptune did turn out to exist, which completely scuppers the notion that "ad hoc" automatically means "wrong". 

In some extreme views, there are no ad hoc hypotheses at all : they can't be clearly defined and depend too much on circumstance, and since they can turn out to be right, there's no point in distinguishing them from any other hypothesis. The authors here note that both proponents of modern dark matter and those of modified gravity view the other as embracing ad hoc hypotheses in a pejorative sense : to dark matter adherents, modified gravity does nothing except explain rotation curves; to modified gravity researchers, dark matter does nothing except... explain rotation curves. And such hypotheses can be both conservative (seeking to save existing ideas, like Neptune in a Newtonian framework) and progressive (like Zwicky's dark matter in an EdS universe).

I think my take remains that the most important thing for a hypotheses is testability rather than how many ideas it explains. True, we might get a bit suspicious if an idea is invoked to explain a single, unique observable, but this is all we should do, rather than insisting the idea was no good. If you have no other grounds to suspect something, invent it on the fly to explain just one thing, and don't have any reason to expect you'll be able to use your idea elsewhere... then your idea might just be of low overall importance rather than actually wrong. It may, in fact, be perfectly reasonable to suspect the existence of a particular planet or galaxy based on observational evidence, and this may be of locally extreme importance : it just isn't likely to alter anything fundamental.

Where the concept of an ad hoc hypothesis does start to become more problematic, I think, is where it is invoked to explain the fundamental basis of a theory. If you need it to explain a single observation, but without this the whole theory collapses, then this should give you pause for thought. By no means does it suggest the hypothesis is wrong, but it's clearly better if your idea explains multiple things or a general situation rather than just one specific thing.

Neither the modern or Zwicky's concept of dark matter constitute such a thing. Both were and are used to explain multiple lines of evidence. In Zwicky's case, some of those aspects were his own ideas but some were completely independent. In that sense, say the authors, Zwicky should be recognised as the "quantifier, not the discoverer, of dark matter". He used methods both of his own and others devising to explain both his own and others observations ; the idea of dark matter itself was not original to Zwicky. Here irony piles atop irony. His findings were correct but unconvincing, with his paper not cited for 25 years; his value agreed with a theoretical framework which turned out to be completely wrong; his basic technique correct but involving a wild extrapolation. 

In this reading, Zwicky comes through as both a revolutionary and a staunch conservative. He had the best of ideas, he had the worse of ideas. But, pretty decisively, it seems we should give up on any misconception he simply invented dark matter to explain a few errant galaxies. Rather, in this particular case, he seemed to have been doing good science, the best he could do at the time – subsequently revised, but that's exactly what any good scientist can hope for. Zwicky did indeed push the boundaries forward, and if others had paid a bit more attention, the history of cosmology could have been completely different.

Thursday, 5 March 2026

Another One Bites The Dust

One of my all-time favourite dark galaxy candidates was discovered by FAST back in 2023. An isolated gas cloud with no obvious optical counterpart, it also had an apparently flat rotation curve. This is the quintessential dark galaxy candidate : small, to be sure, but it would be simply unrealistic to hope for more.

Two small concerns stopped it just short of the Platonic ideal that would convert skeptics into believers. First, its flat rotation curve was marginal, because it wasn't well-resolved by FAST. The 500m telescope is enormous, but to see structure properly, you need something like the multiple-kilometre baselines of the VLA. Second, the optical data of the area wasn't especially deep, and it was always possible that a faint optical counterpart might be lurking there.

Earlier this year came two papers in quick succession, both targeting this same object. The first used deep optical imaging to search for faint stellar emission, while the second used the VLA to better resolve the HI detection itself. Both came to the same conclusion. This isn't a dark galaxy, it's just an unusually faint object.

The first paper is straightforward enough. They pointed a relatively small 1.4m telescope at the gas cloud for a whopping 27 hours to get really deep data, and lo and behold, a galaxy appears ! A small, faint one, to be sure, but it's very clear. It's a little bit offset from the position determined by FAST, but this is entirely consistent with FAST's big beam. And then, just for good measure, they observed this for three hours on a much larger 6m telescope to get the optical redshift... and bingo, it matches that of the HI cloud. So this is definitely the optical counterpart, no ifs or buts.

The odd thing is that the stellar mass of this object is much higher than the upper limit FAST determined could evade detection. Based on the discovery paper, I said that this would almost certainly turn out to be something really interesting, but the stellar mass of the counterpart turns out to be almost exactly ten times greater than what FAST predicted could be there. This makes it unusually faint, but by no means exceptionally so. The authors don't care to venture a guess as to why the upper limit from FAST was so much more optimistic.

The second paper arrives at the same conclusion by a different method. The better resolution of the VLA means they can more accurately locate the centre of the gas cloud. They don't have the same deep optical data as the first team, but by carefully stacking PanSTARRS data, they can do well enough. And lo, the same galaxy emerges from the noise. They estimate the stellar mass to be about half that of the first team, but factor two variations in stellar masses are quite normal at the best of times – and with the noisier optical data this group have to work with, this isn't surprising at all. It's still much more massive than the original upper limit.

The VLA data doesn't show the same flat rotation curve as the FAST data did, but it does clearly show that the gas has ordered motions. That's actually the more relevant factor, demonstrating that it is indeed rotating and so likely stable. So again, it's definitely a galaxy, and a pretty normal one at that.

To be fair it is notably faint. The second team are open to a bit more speculation than the first, suggesting that this might be something similar to the notorious Ultra Diffuse Galaxies but a bit smaller. And it's true that while being a bona fide dark galaxy would have been way cooler, it's still important to study objects like this : they still raise the question of what keep star formation suppressed, albeit not quite so dramatically as originally suggested.

Nor does this rule out the notion of "dark galaxies" more generally. What I think we can rule out, pretty definitively, is the idea of massive dark galaxies. There may very well be absolutely no truly dark galaxies as massive as our own Milky Way (although for sure there may be some very faint ones), or even a fraction of our mass. Okay, sure, in all the unimaginable vastness of the cosmos, there might be one or two hiding somewhere. But as a population, this is something I think we can now safely rule out.

What we can't and shouldn't rule out is the idea of dark minihalos. The reason dark galaxies originally gained traction was to explain the "missing satellite" problem of the Milky Way, where the very smallest dark matter halos never accumulate enough gas to form stars. This would explain the huge discrepancy between theory and observations – and these kinds of dark galaxies are still very much permitted by the data. Indeed, it was only a few theoretical models that allowed for massive dark galaxies at all, so dark satellites remain as plausible now as they did 30 years ago. 

Finding those, however, is turning out to be bloody difficult. The challenge, I think, lies mainly on the theoretical end, in proving whether we really expect any of them to be detectable with radio telescopes at all – and if not, in establishing some other way of verifying their existence. At the moment they feel a little bit unfalsifiable : not to a fatal extent by any means, but to a degree where we should definitely be going "hmm" and doing a good bit of head-scratching. Nothing wrong with that... it wouldn't be research if you knew what you were doing.

Wednesday, 25 February 2026

An Assemblage of Fluff

I'm happy to report that the refereeing process for my latest paper was one of those instances where I couldn't honestly find anything to complain about. In fact the reviewer provided several papers to cite which were, unusually, all genuinely relevant and interesting. And they weren't even an obvious case of "please cite me !" as so often happens with suggested citations ! Goodness, a reviewer actually providing interesting things from a desire to be helpful, whatever next...

Anyway, if you don't want to read my paper or the linked blog post you don't have to. It's basically a very simple one : we found an Ultra Diffuse Galaxy apparently losing gas in the Virgo Cluster. You might remember another such claim (which I've covered a few times previously, of which this post summarises the current state of affairs) which has proven controversial. This one, I think, should be relatively unambiguous as far as the main claim goes. What will likely cause more of a discussion – if anyone notices it at all – is the assertion that this might, like other UDGs, have a depleted dark matter content.

But what I want to do here is give a bit of a speed summary of the other papers I ended up reading as part of this. The referee seems to be of a background from the world of globular clusters, which is far outside my area of expertise. I tend to not pay enough attention to the studies of globular clusters in UDGs, which is a bit silly because that's a whole other area in which UDGs have proven controversial.

One of the claims in my latest paper was that if our UDG was typical of other UDGs in clusters, this implies that many of those also lack dark matter. The referee pointed out that there are different claims regarding the globular cluster content of field and (galaxy) cluster UDGs, meaning that they likely don't have the same origin... which rather limits the impact of what we can state regarding this one object. This led me on a paper chase down a rabbit hole, and for the sake of not having to write out everyone's names the while time, here are the links and abbreviations :

Forbes+2020 : F20; Lim+2020 : L20; Benavides+2021 : B21; Grishin+2021 : G21; Junais+2022 : JA22; Jones+2023 : JO23; Hartke+2025 : H25;  Sandoval Ascencio+2025 : SA25. 

Right then, with no less than eight different papers to cover, let's do this one thematically. There are a number of issues which most of the authors addressed in different ways, so I'll try and unify what they were all getting at.


Where do they become diffuse ? 

That is, are they born big and fluffy or do they have big fluffiness thrust upon them ? Opinions differ, but here the consensus is reasonably clear. B21 argue that some field UDGs are too isolated for environment to have had much of an effect so they must be intrinsically large objects, saying that infall into a cluster will change other aspects but not size. F20, L20, and JO23 all imply that UDGs must have been large before cluster infall, but only on the circumstantial grounds that the differences between field and cluster UDGs are too stark to allow the one to evolve into the other. Only G21 argues directly against this, saying that modelling shows the removal of gas from UDGs significantly changes their gravitational potential, thus allowing stellar orbits to considerably expand.

I'll go through some of the more detailed arguments below, but I generally agree with the majority here. There are more than enough UDGs detected in the field that there simply doesn't seem to be any real need to invoke environment to explain the size of UDGs. It's entirely possible that some UDGs which now seem isolated lonely hermits were once denizens of more hedonistic environments, but it's not credible that this is true for all or even most of them. 

I do think both B21 and G21 might be overstating things though. Removal of gas certainly does have the capability of making some galaxies more diffuse, but this alone can't possibly explain the majority of isolated UDGs. In short, environment is likely only a contributing factor, and not the dominant one.


Where are they quenched ?

Closely related to where they become diffuse is the question of where they stop forming stars. Broadly, most cluster UDGs are red and dead, whereas many field UDGs are blue, gas rich, and star forming. But there's a big selection effect here : cluster UDGs are easy to spot whereas field ones are much harder. The issue is distance. It's perfectly reasonable to assume that a large population of galaxies which is aligned on the sky with a cluster, and not found in control fields, is at the same distance as that cluster. In the field you don't have any such corroboration. Gas measurements can provide distances, but that leads to the selection effect : if you select by gas, of course you'll find that field UDGs are blue, gas-rich party animals.

Still, overall this picture holds up very well. Hardly any cluster members have been detected with gas, which is not unexpected : it makes complete sense than an infalling UDG will lose its gas and stop forming stars – in fact, it can hardly avoid it. The issue is more whether this can also happen in the field rather than whether it happens in clusters at all. While B21 say that even apparently very isolated, quenched UDGs are probably "backsplash" galaxies (that is, cluster members that escaped), SA25 say they've found at least two UDGs where this is unlikely.

It isn't clear to me why SA25's field-quenched galaxies can't be the same "used to be cool" UDGs that once hung around in rowdy galaxy clusters, as B21 prefer for similar objects. But it's crucial here that B21's conclusions are from simulations, not observations. How, then, could we tell them apart observationally ? How could we say that an isolated UDG was once quenched in a cluster versus in situ ? This is far from obvious. 

Clearly, clusters can and do quench UDGs, but whether they can also quench in the field is not at all settled. The potential mechanism by which this would happen turns out to have very interesting consequences, but we need to examine a few other points before we get back to this.


How do field and cluster globular cluster populations in UDGs differ ?

Globular clusters are a component of galaxies I probably don't know enough about. They're dense little starballs that orbit throughout a galaxy's halo, and as JO23 describe, they're thought to form at the very first stages when a galaxy assembles. Establishing how many GCs a UDG has, though, seems a matter of some difficulty. Most studies appear to do this statistically : they use some selection criteria to see how many GC-like objects they find in a control field, then count how many similar objects they find in close proximity to a UDG. This means that in some cases, the estimated GC count for a galaxy can actually be negative since it might have less GCs than the control field predicts.

The situation is not at all clear. F20 found significant variations in the UDG-GCs within the Coma cluster, with some being especially rich compared to comparable non-cluster galaxies, while others were more typical of the populations seen in dwarf galaxies. Fornax dwarfs appear to have fewer GCs per galaxy, while L20 find that Virgo UDGs have more GCs than Fornax but less than in Coma. 

JO23 go on to make a bold assertion that since gas-rich field UDGs have less GCs than in clusters, this definitively rules out field infall as the major origin of cluster UDGs. To be fair, it's very difficult to see how an infalling UDG could gain globular clusters, as long as the standard paradigm that GCs only form early in the universe is indeed correct. This is a pretty strong argument, but I'd add two major caveats : first, we know star formation can be triggered by the onset of ram pressure stripping, so maybe GC formation is also possible; secondly, less speculatively, I'd like to have much more secure data on the cluster GC population. Pretty much everyone seems to agree on a bimodality of cluster GC populations, with "multiple pathways to UDG formation" being an almost ubiquitous phrase in all of these papers.


What is their total mass ?

This bimodality is strongly evident when it comes to their mass estimates as well. F20, L20, J22, JO23, and SA25 all suggest that at least some UDGs are basically normal galaxies : the extreme tail-end of a population that happens to have a very low surface brightness, but with the same formation mechanism at work. Some of these authors (more than I was expecting) quite strongly support the notion that some UDGs are likely "failed giants", with a total mass comparable to the Milky Way, while allowing other UDGs to be essentially just faint dwarfs. 

JO23 expresses this last option most bullishly, dismissing the interesting claims – and I think here very unfairly – of UDGs lacking in dark matter as being problems of data quality. In this view there's nothing much to explain, since most UDGs would be entirely normal objects, just the extreme end of those with typical characteristics. This is an especially weird and surprising perspective to me, since Jones has been at the forefront of research into so-called "Blue Blobs" : stellar structures which resemble galaxies but appear to lack dark matter !

F20 and others express in my view a more reasonable opinion. Some UDGs are quite likely just extreme dwarfs while others are failed giants, and everyone accepts that multiple formation mechanisms are at work to some extent. It was very surprising to me that only JO23 mentioned the possibility of a dark matter deficit, and then only in critique – but this may reflect the GC-centric bias of these studies.  Because they're relatively easy to detect, GCs have become popular as a proxy to estimate the total mass of UDGs, and this is far easier than trying to measure the dynamics directly. It feels like the overall mood of the community here is geared towards measuring dark matter excess through GCs rather than deficit, which it itself quite interesting. But as L20 point out, this requires extrapolations across four orders of magnitude so maybe this isn't such a good idea.


What effect does environment have ?

J22 run some detailed modelling to see if ram pressure can reproduce the observed properties of cluster UDGs. They find that it can, which isn't that surprising : as you'd expect, it makes them red and dead. Furthermore they find that without this gas-loss based quenching, they can't get results which agree with observations. It seems like ram pressure stripping is very much a key aspect of the evolution of cluster-member UDGs. But JO23, again I think rather unfairly*, criticises this on the grounds of the GC population. The thing is that J22's study was based in Virgo, where from L20 we know that the GC population isn't much like those of other clusters, and we also know that GCs vary considerably even in the same regions. 

* Apologies to Jones, but while I've often told anyone who'll listen about the importance of BBs, I can't say I agree with this particular work all that much.

It's also worth noting that the L20 UDGs agree with J22's scenario quite nicely, being generally redder and more numerous towards the cluster centre, but L20 themselves note that this is in contrast to previous findings. So there's a lot of controversy and complications here, and I wonder if there's a degree of mixing apples with oranges. Sure, not all GC changes can be explained as the result of field infall, but that doesn't invalidate the result that ram pressure is needed to explain other UDG changes – and those GC differences might not apply in Virgo anyway. Oh, and the L20 sample is rather small, just for good measure.

On another front, H25 present the case of a very interesting individual object which might just poke JO23 gently in the eye. They describe a UDG aligned with a stripped ionised gas filament seen from a larger cluster member galaxy, speculating that we're seeing a UDG actually in the process of formation. This object has no GCs of its own, which would definitely be at odds with F20/JO23's claims that cluster UDGs have more GCs than field UDGs. 

Now that's potentially a really cool and weird result, but unfortunately this object is (to my mind)... totally unconvincing. It's perfectly smooth and structureless, with a considerable offset from the gas filament rather than a neat alignment. It's also red, implying it's already been stripped and quenched, but this makes very little sense if it's still in formation or was formed very recently. They suggest dust could explain this, but dust is seldom detected in low surface brightness objects (if it was external dust, this should be directly measurable by how much other objects in the vicinity appear reddened). The age estimate of several gigayears is also at odds with a recent formation, which ought to suggest an age of < 1 Gyr. 

In my view this is no more than a projection effect, a perfectly innocent gas filament happening to appear near the UDG on the sky but without any true association between the two.



What I mainly learned from all of this is that UDGs are complicated.

The overwhelming consensus is that there are many different ways to skin a cat : almost certainly, some UDGs form by different mechanisms than others. The issue is where the balance lies, in establishing if any one scenario is more prevalent than the alternatives, and trying to find some common factors to give us a unified picture of what's going on.

Certainly I think the argument that field and cluster UDGs have different GC populations is a very powerful and important one. Were it the other way around, the situation would be different. It's relatively easy to imagine how a GC-rich field UDG could lose its GCs in a cluster, but gaining them ? Maybe not impossible, but it's hard to see that happening much.

The final interesting point is that both SA25 and JO23 propose a mechanism by which UDGs could be fluffed up even in isolation. They say that since they're extremely diffuse objects, they might experience bursty star formation rather than the more usual stable, continuous variety that occurs in brighter objects. This would periodically redistribute the gas, keeping it at a relatively low density. Every so often it would recollapse to a denser level and trigger another burst of star formation, and the cycle would repeat, eventually blowing out all their gas permanently and quenching them completely. UDGs might even then be a normal phase of evolution in dwarf galaxies : a temporary, transient state rather than a genuine type of object.

Intriguing... but I'm skeptical. If this was the case then I'd have expected simulations to have predicted this well ahead of time, along with the results of an apparent deficit of dark matter. We also don't see especially extended gas discs in field UDGs, nor (to my knowledge) any particular offset between the gas and stars. What few resolved gas observations we do have look like normal, stable, circular-ish discs. 

All this underlines just how far we have to go. We have a great deal of data in some areas, but in others we have a woeful lack of even the basics. We don't have a common framework to decide when a UDG might be a failed giant or a dark matter-deficient object; weirdly, these two very different circumstances appear to result in objects which are optically indistinguishable. The old cliché is in this case emphatically true : more research is needed.

Thursday, 18 September 2025

Weaponising dark matter

Stephen Baxter's Xeelee sequence revolves around a war between baryonic and non-baryonic life forms. One memorable sequence features a pulsar being hurled at the Great Attractor (because reasons). Today's paper feels like it could easily fit within such a realm of the gloriously far-fetched, albeit it's not without some reasonable evidence too.

This is just a four-page letter so I'll keep this one very short indeed. Like the last paper, they claim to have found the signature of a very small dark matter halo, but this one's even smaller. The last one was about a billion solar masses, extremely small by galaxian standards but not outrageously so... this one, by contrast, is probably no more than a few tens of millions of solar masses, with a lower limit of just a few thousand.

Such features are certainly predicted in cosmological simulations. Basically, the higher the resolution, the more small dark halos result. But below a certain limit, nobody ever expected to have much chance of ever detecting them, since they'd have so little gravity they'd never attract enough gas to form a single star. And once you add in all the baryonic matter to the simulations (the boring normal matter of stars and gas), presumably most of the smallest ones would be disrupted.

The claim here is they've found a bullet wound in the Milky Way resulting from a collision with one of these minihalos. Actually, again like the last paper, this is not a discovery announcement so much as an independent confirmation by a different method. The original discovery came back in 2017 in the form of a molecular gas cloud with an unexpectedly high line width. 30 km/s is small by the standards of galaxies (the Milky Way would be more like 400 km/s), but with no stars to drive the motion, dynamics have an obvious appeal. Without any other visible material (at least nowhere near enough), dark matter is at least heavily implied.

Here they use Gaia data to look at the velocity of the stars in the vicinity and discover they have a vertical velocity anomaly : in a small region of the disc, the average velocity of the stars perpendicular to the disc drops, even while their dispersion increases. The original CO blob is slap-bang in the middle of this VVA, with the VVA being very much larger than the CO blob. Which would be an awfully suspicious coincidence.

A lump of dark matter colliding with the disc could certainly cause this. Its small size is certainly consistent with a very small halo. But I have no familiarity with stellar dynamics on these scales at all, so I can't tell you how unusual such features are and their figures don't really give much of an indication. They also don't even consider other explanations, which I suppose is fair given the limited available space, but it would have been nice to have mentioned something (besides the obvious impossibility of stellar winds and the like, there apparently being no stars here). And how much such features would we expect to find, if these minihalos exist in the numbers predicted by simulations ? Finally, their tie-in to Ultra Compact Galaxies more generally – which are largely thought to be the stripped cores of more massive galaxies – is just too speculative even for a letter.

In short, it's definitely a very interesting feature to report, but it's going to take a lot more work to say anything definitive about what it actually is

Monday, 15 September 2025

A dark RELHIC of an earlier age

The last time I tried to count the number of times objects had been claimed to be the first dark galaxy candidates, I stopped at ten because I got bored. Today's paper adds another one to the list.

To be fair, not all dark galaxy claims are equal. Some would say a galaxy only counts as dark if it really has no stars at all, others that it just needs to be sufficiently dominated by its gas and/or dark matter. Others would insist that it has to have certain dynamics or only be found in the nearby universe*. Most would probably demand it had a primordial origin rather than just being stripped out of a galaxy, but not everyone would agree. So that list of ten could very plausibly be extended or contracted considerably.

* Pretty much everyone agrees that galaxies started off as dark, so we accept that dark galaxies did exist at one point. The controversy is over whether any still remain dark today.

This paper concerns a very particular type of dark galaxy they annoyingly dub a RELHIC. Why annoying ? Because we also have radio relics, which are completely different beasts : they're incomparably larger and more diffuse, and have little or no direct relation to individual galaxies. These objects, on the other hand, are Reionisation Limited HI Clouds, a term coined by one of the authors that I'd urge them to stop using. 

But no matter. What they report on is a very interesting object that in some ways is the sort of dark galaxy candidate everyone wants to find. One of the major reasons to suppose such objects exist at all is that they would solve the long-standing missing satellite problem, whereby simulations produce far more galaxies than are actually observed. The idea is that while the physics of gravity is pretty simple, the physics of star formation is anything but. So maybe these objects do exist, it's just that they've never formed stars. This is now the widely-accepted explanation for missing satellites – indeed, arguably there isn't such a problem at all any more, as (in some models) the only such "dark galaxies" are now so small and so lacking in stars that we wouldn't detect them.

One of the complexities of the physics behind these is that of the Epoch of Reionisation. The first stars as thought to have been super-powerful monsters powerful enough to ionise most of the gas in the early Universe, heating it to the point where it would be driven out of the smallest galaxies completely. Models show that below a critical mass threshold, galaxies would lose all of their gas and never form any stars at all. It's not quite a sharp cut-off, with some near or slightly above the threshold able to form some stars before reionisation brought the process to a permanent halt, but it's close.

Such objects are thought to be extremely difficult to find. Their HI masses should be only a few million times the mass of the sun, about a hundred times less than a typical dwarf galaxy. And their line widths might be only 20 km/s or less, barely wider than the HI line itself . In principle these objects could be almost numberless, just bloody hard to spot. In contrast, most dark galaxy candidates that hit the headlines are much bigger, and usually by the author's own admissions fairly exceptional – massive dark hulks that are relatively easy to find despite being so rare as to indicate little or nothing about how most galaxies form.

Here the authors present a candidate discovered with China's mighty FAST telescope. In contrast to this awful RELHIC term, I can't fault them for the name of this particular object : Cloud 9. Yes, really. Apparently this was first reported in 2023 but I seem to have missed that paper when it came out. 

Here they report on deep Hubble images and confirm that it's really, really dark, with no more than a few thousand solar masses of stars against it's million or so of gas. That, together with its line width of just 12 km/s (!) and small size (1.4 kpc radius), with an estimated halo mass that's extraordinarily close to the mass threshold, make it a compelling RELHIC candidate. It's certainly one of the darkest objects ever found, which is always a pretty cool thing to find.

But just how good, exactly ? My verdict would be... yeah, this one's pretty interesting. Is it definite ? By no means. But it's a good candidate, and absolutely needed to be published.

Glancing at the discovery paper, it seems that Cloud 9 is a little over 100 kpc from M94 itself, with HI clouds closer to the galaxy that are clearly some form of debris. 100 kpc is quite far, but only a few times the size of a large galaxy, and certainly there are many extended streams known which are much larger than this. So this cloud could be a leftover far-flung bit of debris as well, but like Cloud 6 in Leo, it doesn't really fit the general pattern of the other clouds. 

A perhaps more serious difficulty would be that estimating the total mass of the feature must be extraordinarily difficult. HI tends to be found at ~10,000 K, corresponding to a line width of 10 km/s. A width of 12 km/s tells you pretty much nothing beyond the thermal state of the gas, so inferring its dynamics from this is... well, my worry is that you simply can't get a meaningful estimate when things get this narrow. Even if there was no dark matter here at all, the line width wouldn't get much narrower because this is about as narrow as the line can get. It's not a matter of better data in this case : nothing will help, at least not very much.

Another issue is the question of how long the cloud could survive, and conversely, how long it's been in existence. Currently its gas density is well below the threshold for star formation. Assuming it began life so small that the density would reach the threshold (so as to have always been dark), and given its expansion velocity and small size, it would have taken perhaps a hundred million years to reach its present size (without dark matter, as expected if it's just debris). In a another hundred megayears or so it'll double its and likely be undetectable. So if it's debris, we're detecting it at an unusual point in its existence, but sadly this doesn't constrain things too much.

I think this is a case where what's needed is a good set of simulations, especially given the timing constraints from the size of the cloud. What kind of interactions could affect M94 that would produce debris like this ? How often are such things formed, and are those simulations compatible with all the other data of the system ? What happens to existing minihalo RELHICS like this one in a system like this, where there's clear evidence that M94 experienced a merger – can they survive in such a place ?

This one's going to take a lot of work to answer. It would be easy to dismiss this as just another bit of HI fluff... and it might be. But it's so close to what we expect minihalos to be like that the workload might be worth it. And perhaps, just maybe, some other clouds already lurking in the data aren't the boring bits of debris we all thought they were.

Friday, 12 September 2025

The galaxies that seemed magical are actually just very lazy

Today, two papers for the price of one ! The one is a sequel to the other and they're both quite technical, so let's knock off two birds with one stone in the bush and other mixed metaphors. Paper I is here and paper II is here.

The papers in essence address two related questions on Ultra Diffuse Galaxies – the big faint fuzzy things that often seem to lack dark matter, which I've covered here ad nauseum. Neither paper much addresses the dynamics (i.e. total mass) of the objects, but rather the other fun aspect of these galaxies : why are they so wretchedly bad at forming stars ? Many of them have tonnes of gas, so why aren't they forming stars like normal galaxies are ? Why are they so large and yet so faint ?

The first paper deals with the basics. It tries to determine if the star formation efficiency of UDGs really is weirdly low, or if this is only a selection or measurement effect. Spoiler alert : it really is low. So the second paper then address why this might be – is there something missing in our basic model of star formation, or is it just due to the peculiars and particulars of of these particularly peculiar systems ?


Paper I begins by collecting a sample of 22 UDGs and 35 more typical dwarfs of comparable mass.  The UDGs all have atomic HI gas detections, though low resolution so essentially all we know is the mass of gas, nothing about its structure. But it seems that at first glance, UDGs are indeed of systematically lower star formation efficiencies : their star formation rate is less than that of other galaxies of similar gas masses, and given their stellar masses they have more gas than expected. Both of these effects are modest though. It's quite apparent that the population as a whole is systematically offset from the rest, but they're all still within the general scatter. Interestingly, they also note that the properties of the HUDGs (HI-detected UDGs) aren't much affected by environment.

The first question they tackle is whether these objects are really different in terms of their gas content, or if this is just a selection effect. That is, the HI observations might be limiting what can be detected at all. It could be that those of lower gas fractions do exist, it's just that the data isn't sensitive enough to show them. But they find that's not the case : they should be able to find considerably less gassy-objects, so these don't seem to exist at all. These HUDGs are not the tip of a less-gassy iceberg.

Next, their main topic. While the average star formation efficiency of HUDGs appears low, this could just be due to the statistics from using the total gas mass and stellar mass, which smooths over the whole structure of the galaxy. It could be that the star formation efficiency is actually quite normal, just restricted in area. For example the gas could all be concentrated in the centre and forming stars pretty normally, but when averaged over the whole galaxy, this would be "washed out" and it would look like the galaxy was rubbish at forming stars. Overall, they would be, but locally, they wouldn't be that bad. This would imply a significant population of older stars outside the gas-dominated regions.

To test this they use spectral energy distribution fitting. Basically what this does is use many different data points across the optical, UV and IR spectrum to estimate the stellar ages as accurately as possible – this is the best we can do in terms of estimating a galaxy's star formation history. Ideally we'd also like to have resolved measurements of the HI, which sadly they don't have here. But they do SED-fitting for many different points per galaxy, so they can see if the low SFE is something that varies throughout each object or if it's low everywhere.

I'll skip over the details of the SED fitting because I don't understand any of it; I only note that they stress they aren't constructing detailed star formation histories here, just enough to answer their main questions. Their main result is that the low SFE is true on all scales from big to small. It's not just that the gas could be more extended, it's that UDGs are bad at forming stars full stop, even if the gas density is higher. While UDGs are large given their stellar masses, they're not especially large considering their HI masses.

There are a lot of different scaling relations to juggle here, but the end result is very simple : UDGs aren't good at converting gas to stars even when they've got plenty of it. The obvious next question is, of course, why are they so bad at this ? For this we need the sequel.


Paper II takes quite a different approach and is all about modelling. There's a really popular and widely-used relation between the surface density of gas and its star formation activity, but there have been indications for many years that we ought to be using the true, volumetric (3D) density instead. This is harder to measure directly so some assumptions have to be made, but it can be done. 

Here they consider a particular version of the volumetric star formation law that depends on the different components of a galaxy together : the gravity from the atomic and molecular gas, the stars, and the dark matter. All have different gravitational contributions. For instance in the centre of a galaxy everything might be extremely dense, whereas further out their might still be lots of atomic gas and stars but less molecular gas, and on the very outskirts only atomic gas and dark matter. So even if the atomic gas has the same surface density, it may need the extra gravity of the stars to help pull it together and collapse.

Again I shall spare you the technical details of the model. This time my essential note is that they consider UDGs to be rather more dark matter dominated than ordinary galaxies, which flies against the prevailing winds in the last few years. More on that in a moment.

They find that this more complex model... works ! It can explain the star formation rates of both normal galaxies and UDGs under very reasonable assumptions : there's no need for any additional physics, no weird mechanism or alien interference that suppresses the star formation. There are many uncertainties, but their sensible default assumptions are enough to give a good result without any sort of fine-tuning needed. So UDGs are, in a sense, pretty normal.

They also find that the model isn't very sensitive to how much molecular gas the UDGs are assumed to contain. That's bad news for anyone trying to detect their molecular content : essentially it implies they could well have very little of it. They can even estimate just how much molecular gas they expect them to have, given their estimated star formation rates and the (rather surprisingly but repeatedly established) independent finding that molecular gas generally has a constant depletion timescale of about 2 Gyr. In short, their conclusion is that detecting the molecular component will be Bloody Difficult. Dwarf galaxies are already hard to detect, but UDGs will be even worse.

What about that choice of a relatively massive dark matter halo ? They explore this, and rather surprisingly it doesn't matter much. It seems that the density of the dark matter is anyway assumed to be low so that reducing the total mass doesn't make a great deal of difference. In fact this can give slightly better agreement with the measured star formation activity, but unfortunately, they say there are too many other uncertainties to say if their model prefers no dark matter at all to a normal mass halo.


So, there we have it. There's no need for any weird physics here : star formation in UDGs apparently follows the same laws as in every other galaxy. The difference isn't so much the gas content itself as the state it happens to find itself in, just as a puppy can be an energetic ball with the madness of a thousand caffeinated suns or the sleepiest thing since Slothy the Sleepy Sloth swallowed an entire bottle of Nytol after a mug of warm milk in a comfy armchair. 

Does this mean UDGs aren't weird at all though ? Nope ! You might remember that there were previous claims that UDGs are actually just normal galaxies if you redefine how to measure their radius. That was true, but doesn't mean that the other radius estimates were wrong : it still points to UDGs being anomalous, just not in the way we'd understood. 

Here we still don't know what sets the initial conditions of UDGs. Why do they start out so differently to normal galaxies ? Is it their dark matter halos (or lack thereof) and if so, is this compatible with our simulations of the sort of halos we expect to actually exist ? And how come their globular cluster populations appear to be very different to other galaxies ?

We still don't know. We do know that UDGs are weird, but as to whether they're pointing to a deep flaw in our models or just some incompleteness or other... the jury's out. Or more accurately, the trial hasn't even started yet : we need more evidence before we can even really begin. But at least what this paper adds is what evidence we should go after. If UDGs are found to have chonks of molecular gas, that would falsify their model straight away. If they're not, we keep investigating.

Why Bother ?

It's rare that I manage to read any longer pieces on arXiv that aren't strictly about galaxy evolution, but today I indulge myself. ...