Sister blog of Physicists of the Caribbean. Shorter, more focused posts specialising in astronomy and data visualisation.
Showing posts with label Philosophy of Science. Show all posts
Showing posts with label Philosophy of Science. Show all posts

Monday, 27 July 2026

Why Bother ?

It's rare that I manage to read any longer pieces on arXiv that aren't strictly about galaxy evolution, but today I indulge myself. And I'm glad I did, because this particular piece was highly provocative and well worth reading in full. This write-up will itself constitute something of a long read, so I advise getting some tea before we begin.

Ready ? Good. Here goes then.

David Hogg's self-proclaimed "very white" paper thankfully isn't so in the horribly racist sense, but in the "here are some semi-organised thoughts that might be worth considering" sense. It's entitled "Why do we study astrophysics ?" and it's ostensibly written in the context of ever-increasing AI development. His goal is to provide a stepping stone to understanding when and how LLMs should be employed for astrophysical studies. He doesn't attempt to come up with a final answer to that question, because conscious or not, the effect of a machine that can answer questions more accurately than human experts still represents a technological and sociological singularity. Seeing beyond that point is too big of an ask. 

Instead, he tries to tackle the fundamentals of why we do our job at all, thinking that this is a necessary precondition for how and why we might go about automating parts of it. If we can figure out why we're doing what we do, maybe we can better understand what we should do given expected technological progress.

As an essay, I found this one thoroughly excellent. There's much here I disagree with and much I support, all of it well-argued and clearly stated. It sticks to its central theme but covers a very wide array of topics along the way. There's actually not too much about AI in here, the focus being more on the human side, but I think it may have at least a germ of an answer as to how we'll proceed come the killer robot uprising technological singularity.

To be fair, Hogg's essay isn't the most linearly organised piece in the world, often feeling like something of a memoir. I've tried to keep this summary-cum-commentary to the linked themes and bits I thought I had something worth contributing to; I've deliberately avoided issues where I disagree but don't think the argument would get us anywhere (such as whether LLMs are truly thinking or not... this is very interesting to me, but makes no difference to the arguments here). 

First, I'll look at the main thrust of the article : what astrophysics is and why we do it, which involves various thoughts on how we go about this. Then I'll conclude with a much shorter section on what this might mean in that future where we have vastly more powerful analysis capabilities than anything we possess today, but which we can reasonably expect to have in the coming years.


1) Astronomy Today

The professionalism of science 

Much of astronomy is now done using large facilities by enormous groups, and the way these operate is inevitably different to small groups in a lab in someone's basement. Here I think the word he's searching for is industrialisation, not professionalism. The career-based nature of astronomy has already been long established, but the escalation of scale is still relatively new. The point is that you can't have people just casually mucking around on dedicated survey tools or slapping their own instruments onto billion-dollar facilities. This was absolutely possible in the old days – the underside of the Arecibo dish was littered with discarded receivers – but this kind of approach is all but dead already.

Astronomical data production is becoming extremely professionalized, and in a very particular way. Astronomers, in the case of Gaia, are just end users; end users of curated, calibrated data, delivered by a combination of the ([military-built] secret) spacecraft and the (absolutely great, professional, and open) DPAC... it wasn’t built or operated by astronomers.

Even university-based projects endeavour to produce science-ready data products that can be queried through application programming interfaces, plotted, and analysed without much worry about where they came from or how they got here. I have been involved in bringing about this change, and in many ways it is absolutely great. It democratizes astronomy, since it lowers barriers to entry. It creates an open-science space, in which every project benefits from the output of every other project.

But it does have a strange consequence, which is also related to professionalization : for some kinds of projects in astrophysics, there isn’t a huge difference in capability between a classically-trained astronomer and a newly trained data scientist... a data scientist who has taken an astronomy class might be better prepared than an astronomer who has taken a data science class.

Which is relevant, of course, because LLMs can absolutely do data science.

I think this kind of development is only partially inevitable though. We need big data and big data needs big facilities. But we also need small data, that is, data we can analyse in extreme detail. That may still need big instruments but it also needs small groups, and small groups can, and should, continue to operate in their current fashion. Analysis of big statistics is surely going to change, but on smaller scales, perhaps, is a realm where we can still do the low-level stuff ourselves. Here is where we can invest considerable time doing some aspect of the data reduction and analysis by hand and still have a good chance of making useful, interesting discoveries.

My reasoning is purely pragmatic. Manual data reduction and analysis can be extremely time consuming, but is perfectly manageable on small data sets. We've already abandoned this approach for the largest data sets because it's just not possible to do them by hand. But I fervently believe there is a great deal of value in learning to do the low-level stuff (the hard bit), even if later you never do it again. As a general rule, the better you can operate without a specialist tool, the better you'll be when you get to use it. Maybe for humans, then, the future lies not in big data, but in small data, in extreme specialisation rather than generalisation.


What is astrophysics ?

It's that which produces novel information about the Universe, says Hogg. Reading about it doesn't count, you have to do actual research. But... you also have to document the results. Unrecorded data doesn't contribute to the pool of knowledge from which others draw, so if you don't document it, you might as well not bother. 

Astrophysics, then, is the literature, in his view, and it's this act of producing literature after a novel investigation which best describes the process of doing astrophysics.

I think it's hard to dispute this, but he has a couple of other points which might be more controversial. One is that software isn't as important as the results it produces. I actually do agree with this, because he explicitly declares that software should have associated papers. This then makes it just as important for certain metrics as actual science, and I think it's absolutely right that the effort of software development be properly recognised. 

What he means here, I think, is only that software itself is not science. We write software so that we can analyse data, not because it's intrinsically worth having. Software which isn't used is as pointless as data that's not recorded.

He also notes :

Astrophysics, like any science, contains a lot of “implicit knowledge” or folklore about things like how to observe, how to reduce data, how to organize projects, how to visualize data and models, how to read and write, and so on. Much of this never appears in the literature. Is that not also astrophysics ? Yes it is, but it is astrophysics practice. The results of astrophysics — the scientific conclusions and debates — are in the literature.

Yes, but my answer here would be that we should absolutely record as much of this "folklore" as we possibly can. Some sort of journal of astrophysical methods – not describing mathematical procedures, but the really low-level stuff of what to do with the data and how to interpret it – would be valuable, I think. Such papers wouldn't have the lasting value of results papers, but they would make a lot of people's lives a lot easier. Implicit knowledge should be made explicit wherever possible*.

* Though it is ultimately impossible to record literally everything. Some things you simply have to do.

His second more controversial comment, with which I vehemently but provisionally disagree, concerns papers as a metric :

The second comment is that I often hear software (and hardware and engineering-oriented) people say that they “have to” write papers because papers — and the citations that they generate — are “the coin of the realm.” Papers (and the authorships on those papers) and the citations of those papers are not “coin” of anything ! They represent our recording of what happened, what we learned, what we know, and how we know it. Citations deliver provenance, not reward.

I'd love to agree with this but I can't. In terms of career advancement, it's not sensible at all. Like it or not, astronomy as a professional/industrial career does have certain requirements common to all employment. We need to get paid and we need to ensure job security, and we can't and shouldn't ignore this. The era of the gentleman-scholar is long over : I mean, sure, I'd love to give everyone tenure and a sack of money, but until we do that, papers absolutely and undeniably are the coin of the realm.

In fact, even in terms of strict science, I still don't think I can agree. If astrophysics is the literature, as Hogg claims, and our goal is doing science... surely it's literally true by definition that papers are the "coin of the realm" : inasmuch that if we should be judged by what we produce at all, this should be the primary means by which we do so. I can't get my head around the alternative, which would be a bit like saying that we shouldn't judge a painter on either the quality or quantity of the paintings they produce.


People are the ends, not merely the means

What might explain the above difficulty is what to me feels like Hogg's most controversial and complex point, one which I'm allocating three subsections to examining.

People, says Hogg, are what astrophysics is really all about. It's not about the Universe at all (and he's emphatic and explicit on this point, on which more below), it's about enriching ourselves.

When we employ a graduate student to perform some work, it absolutely must be because the graduate student will benefit from that work, not merely because that work needs to get done. I have heard it said, more than once, in research contexts, that an LLM can do some task “better than a graduate student.” That language makes me uncomfortable, because it is taking an extremely instrumental view of graduate students. Are graduate students in our groups and our laboratories and our universities to do work ? Or are they here to learn ? 

We train PhD students not merely to amplify our own research programs, but to create opportunities, and specifically opportunities for them. Every person is a human being, whose personal development is more important than our short-term scientific accomplishments.

I mean... sure, to a point. I think it's the old Platonic point about whether education is about discovery or change, and it can only ever be both. We shouldn't be using graduate students as literal tools to solve the problems we want solving; they are not there to do the boring grunt work that needs doing but we don't want to do ourselves. But at the same time, we should be getting work done. We should be solving problems ! We should be learning about what Nature is, not just endlessly pontificating on what it might be. 

The only real solution here, I think, is to find graduate students who share our interests so that the result is mutual benefit. We should be giving them problems that both advance knowledge and advance their own abilities. Otherwise, we risk running into one of Plato's weirder quotes (Republic, book VII, 503b) :

Then if, by really taking part in astronomy, we’re to make the naturally intelligent part of the soul useful instead of useless, let’s study astronomy by means of problems, as we do geometry, and leave the things in the sky alone.

I do think it's important to realise that the people doing the work are first and foremost people, not tools. The further we can get away from the mentality that productivity is the only result that matters, the better; the more we can suppress the need to be competitive, the more we can suppress the idea that we need to live to work – even when that work is something we really enjoy – the better the situation will be for everyone. 

But two important tangential points crop up here. First, Hogg declares that this means that not citing relevant papers is (ignorance aside) actually unethical. 

It is ethically required that our papers cite the work that is relevant to the work we are doing. You can’t decide not to cite a relevant paper because you don’t like the author, or don’t like the author’s institution, or don’t like their funding sources. In particular, if the literature gets flooded with work of relevance to your research program, you have reading to do, and citing to do.

I object very loudly to this ! And not just because I'd read dozens of papers that damn well should have cited me but didn't. First, pragmatically, the increasingly industrial scale of astronomy literature means that reading every paper on a topic is not a sane choice. As per the last post, reading hundreds of pages of largely-irrelevant text (and astronomical papers tend to be extremely dry, which is not a minor point) is going to result in negative value, not merely slowing things down. So no, exactly for the sake of not treating people like instruments, you absolutely do not have to read and cite absolutely everything. That's mentally destructive, not productive. You don't have a duty of self-destruction.

Secondly, from a moral viewpoint I also disagree. I see nothing at all wrong in deliberately not citing papers where we don't find the results and/or methods credible, and might even object if I was compelled to cite something I didn't believe – and I'd certainly have a very big problem if I was told by a reviewer to make something sound more plausible than I thought was really the case*. This is not to say we should avoid controversy and it certainly doesn't mean only citing the things we agree with. It only means that we don't have to cite the whole history of a research program and give equal weight to every long-discredited idea or failed avenue of inquiry.

* Hogg actually says himself that you can't cite work you don't trust, but seems to think it's obvious that we can trust people and can't trust LLMs.

There's one aspect here with which I do, however, violently agree :

Every scientific paper is written to help all of its writers, and all of its readers, learn and grow, no matter their career stages.

My take here is not about the content of the paper so much as their style. We need to write for each other, as human beings (which Hogg does very well indeed), not as automatons who require total clarity and unambiguity. If you want me to cite more papers, reform the standard requirements for a manuscript. Make them shorter, better organised (results first, then detailed methods) and more readable (allow the occasional joke, stop being anal about contractions and punctuation FFS). But this is a well-worn hobby horse of mine so I'd best not go down that route again today. 

Hogg goes further and suggests that grant funding to hire students to do work might also be unethical ! And again, I cannot agree with that. We're not running a charity and it's not at all wrong to expect productive output (though we absolutely do need to be flexible in our expectations of that output). This is also in stark contrast to his later claim that we need to use our resources efficiently and get correct, rigorous results. We should seek good working conditions, but ultimately we are doing work. The results do matter.


The answers don't matter

Here we come to the heart of the problem. Hogg genuinely believes that the results of our research aren't important. In one of his oddest moments, he says that if we really cared about the results, we wouldn't do astrophysics ourselves but pay other people to do it for us... this is weird every way I look at it. I just don't think that's how people work, because extending that reasoning, nobody would ever do anything for themselves at all. And of course, it's a pretty perfect example of so-called effective altruism, which Hogg calls an "absurdity" ! He's not wrong about that, but my goodness me, the contradiction is as a glaring as glaring can be.

This baffling oddity aside, Hogg's argument is not to say he thinks we should all quit and do something else. His claim that the results don't matter is more specific and more strict than that... he think the investigations are worth doing, that that's where the benefit lies – in improving ourselves – but that what's actually going on in the Universe is of no consequence to what's going on down here. 

That's, err, quite the hot take there. But it deserves more examination.

Hogg has this highly annoying phrase, "clinical value" which he best expresses thus :

I like to say that the sciences have a “left edge” which is about fundamental understanding, and understanding for understanding’s sake. They also mostly have a “right edge” which is about what I like to call “clinical value” but you could call application or use in the world for technologies or policies. 

I claim (and maybe this is a bit controversial) that astronomy has no right edge. That is, there are no useful things in the world that flow from astronomical discoveries and results. I have spent years of my life estimating the comoving volume of the Universe, measuring the local dark-matter density, and finding planets around other stars. No human outcome or pragmatic capability has been affected in the slightest by any of my results. Literally nothing helpful to humanity arises here.

No sir ! No, I won't have it. First, the reasons why we do astrophysics – the whole title of the essay – are to me obvious. I cannot understand people who don't have any interest in understanding the nature of the world in which we live, and for those that do, then understanding the most miniscule corner of it and ignoring all the rest seems like a clear sign of insanity. For me, observational astrophysics is absolutely a fundamental science, more so, I would argue, than theoretical physics : that's just making up a bunch of stuff, which is valuable, but ultimately tells us nothing about reality, at least not with any certainty. It is observation alone which can do that.

Knowledge of what's beyond the sky is not some abstract wishy-washy thing, but essential in understanding the truth of our own existence. Would it not matter if the stars were holes in the curtain or night rather than fusing spheres of hydrogen ? Would it not matter if the nature of reality were that we were actually inside a giant koala rather than an immense vacuum ? I think it would, and in fact it might well form the basis of all our other knowledge.


Astrophysics is useless

Which leads to the final part of this section, and the second, closely-related reason I think we do astrophysics. Hogg claims that it's not for spin-offs and these don't count as the "clinical value" or right edge. With very few possible exceptions such as discovering dangerous asteroids, and in previous eras understanding chronology and navigation, he claims that these aren't the reasons at all :

Nothing in the world of things or people hangs on the precise value of the age of the Universe. Astrophysics may occasionally and accidentally produce something useful. But astrophysics is not done with the goal of obtaining clinical or practical value. A science has a right edge if and only if the associated clinical work actually tests or exercises the specific results of the science. None of astrophysics is justified in these right-edge terms. No astronomer (that I know) is improving the calibration of JWST instruments because they want the US Navy to have a higher kill rate.

No ! Astronomy's right edge is not in "the clinical value [which] lies in its feeding of humanity’s love" or some other airy-fairy thing that Hogg justly raises as failed counter-arguments. It lies in telling us what is not true. It defends us against ignorance, and the price of ignorance can be extraordinarily high. It becomes extremely difficult to maintain that you need to sacrifice people to appease the gods of the sky when you realise that there simply aren't any. Cosmology has direct moral implications : just because we no longer take a direct moralistic approach to cosmology, as was done throughout medieval history and earlier, and as Tolkien did brilliantly in fiction, it doesn't mean that our morality isn't affected by our understanding of cosmology.

Now to be fair, the precise values of different parameters do not always constitute such a hard right edge. It's not obvious how the exact distance of Proxima Centauri or the HI content of the M31 galaxy could have any moral value whatever. But collectively, we need all these incremental findings to get to the good stuff. We need the flies in the ointment to break our understanding and shatter our conceptual frameworks every once in a while. We need things like the perihelion of Mercury to tell us that Newton is wrong and time and space are themselves not at all what we thought they were, the full moral implications of simultaneity breaking still being something we haven't got a handle on. And we only get those results through slow, careful, methodical measurements.

Does it matter that those of us working on astronomy aren't doing so for the hope of such a breakthrough moment ? Does our motivation being purely intellectual simulation invalidate this hard right edge I've suggested ?

No, I don't think so – not at all. It's true that many of us like the pure research side of things, that we do our jobs (in part) precisely because we can avoid having to be responsible for other people. I too like the fact that nobody's daily lives are at all likely to be impacted by the velocity width of a dwarf galaxy I publish deep in a table of a paper that hardly anyone will ever read. But this does not mean the work isn't worth doing for its own sake, that it won't potentially contribute, albeit in a small way, to the revolutions in thinking which will eventually and inevitably follow. And those who are working with more express goals – if there actually is anyone out there calculating the distance to Proxima Centauri purely to refute astrologers or the hope of fortune and glory – well, more power to them, and equally, their motivations aren't invalidated by the pure interest sake that the rest of us pursue.


2) Astronomy Tomorrow

How does all this mean we should prepare for an astronomical future in the age of AI ?

Hogg proposes two extremes, both of which he views as undesirable. One is that we hand over everything to the LLMs and literally let them do everything, or at most, we try and curate their findings to sort the good from the bad. This would seem to be a pointless exercise in which we don't ever really learn anything, we abandon the joy of the process and reduce ourselves to mere instruments. Pretty much nobody wants that. 

See, I think Hogg does have a point that the human element matters : we do astrophysics for our own enrichment and reward, and both the process and the results matter to us. Even if an LLM had such emotional motivations, there would seem to be self-evidently no value whatever in letting them do everything. That'd be like sending someone to go on a rollercoaster on your behalf. Having the experience, not just having casual access to the results, matters.

When we offload that work to LLMs, we are no longer doing astrophysics, we are no longer becoming astrophysicists, and, eventually, we no longer are astrophysicists. The let-them-cook policy, in the end, leads to the death of astrophysics, the end of astrophysics at universities, and the end of astrophysics education. Astrophysics would no longer be by humans, and then it would no longer be for humans.

The second extreme is that we ban LLMs altogether. This Hogg views as bad because LLMs can be genuinely useful, the effort to seek-and-destroy LLM content would be hugely inefficient and wasteful (again, we'd become mere instruments), and telling people how they can and can't do their research self-evidently violates their freedoms.

Hogg's tentative and intriguing suggestion for a middle route is that we treat LLMs as colleagues who aren't part of our own team :

You might ask your colleague for help finding something in the literature, but you wouldn’t ask your non-coauthor colleague to write the introduction of the paper you are writing. You might ask your colleague to help speed up your code, but you wouldn’t ask your colleague to write your code. 

I quite like this. Necessarily, the middle route must be allowing the LLMs do some of the work rather than all or nothing, and the essence of having some simple guidelines for good practice is sensible. 

I don't think it will work out exactly like this though : recently, I've been "vibe coding" a quite elaborate program and I'm convinced this is indeed the way of the future. I've resisted this practise for a long time, but decided that there was one particular code I really wanted to exist that I didn't have the time to work on (to be shared in a future post). Vibe coding is not zero effort, far from it, but it does work. I very much doubt that we're going to insist on coding by hand for too much longer, any more than we insist on adding and dividing using pencil and paper. Still, the basics of Hogg's idea are interesting.

I need to finish with a few assorted caveats :

Another idea is that, given our respect for our readers, we shouldn’t ask them to engage with something that took way less time to write than to read.

I don't think so. The content isn't any the less valuable based on effort. What seems obvious to one person can be profound to another... I mean, I heard a podcast of Mary Beard – Mary BEARD, for crying out loud – dismissing Marcus Aurelius' Meditations as of "no value". So much for that.

A second is that of the nature of LLMs : according to Hogg, we cannot trust their output, the text isn't meaningful until a human reads it, LLMs aren't reproducible, they currently only produce slop, and they can't take responsibility. All of these I think are only partial truths : we can apply the same methods of trust as we do for humans; the fact that humans will always have to read the text would seem to make the argument that LLMs lack any true understanding to be largely pointless; I agree that LLMs are not deterministic but this isn't quite the same as lacking reproducibility; the idea they only make slop is simply a garbage claim; and much more development is needed here as to what we actually mean by "taking responsibility".

 And finally one throwaway, tangential claim I cannot responsibly let slide : 

Of course it is important to remember that the human practice of astrophysics, at least in its current form, is also very damaging to the environment.

Without reading the citations provided I instinctively think this statement is of negative value. There's no way that astrophysics represents any sort of significant environmental problem. The kind of practises which are truly damaging are those of big businesses, industries, and the exploits of billionaires. It's right and proper that we try and set an example and constrain our environmental footprint. But we should do so only insofar as this helps curtail the much worse damage that's being done by others. Otherwise we risk falling into a Calvinist sort of pointless guilt, in which all we accomplish is to feel horribly depressed for expending energy on a Zoom call while billionaires continue to fly first class across the Atlantic for weekly holidays. 


Conclusions

Phew ! Well done if you made it this far.

I wanted to try and venture a few thoughts as to what might happen next. But, having rewritten this section several times and always ending up with content I never quite believed, I decided to abandon this approach. Instead, I can maybe offer some thoughts about why these predictions are so difficult – and maybe just a hint of something more.

The obvious reasons are that trying to predict what something more capable than ourselves might achieve is fundamentally difficult, and of course the rapid pace of development makes prediction inherently uncertain. A slightly more interesting factor is that different people will adopt different approaches. Some will despise AI, some will love it, some will see it as a tool, others will use it more like collaborators. That we already see this happening makes giving any one answer about what's going to happen next a flawed question, like insisting that there can be only one answer to "what happened in Britain after the Romans left ?" when in fact there are many.

Related to this, I also think that people have very different ideas as to which part of the analysis they'd like to automate away. For me, visual inspection of the data is the fun part. For others, that's the bit they most want to avoid and they want to concentrate on the mathematics or the hypothesis-generation. So prediction difficulties are hit by a double-whammy (at least !) on inhomogeneities : people have different attitudes to AI and different preferences to what they want to automate. Maybe it'll all just balance out.

Another difficulty is feedback : that we don't know how this level of automation will affect us. I'd like to think that it'll be linear, that we all pull back on the stuff we don't like (different though that will be for everyone) and concentrate our mental resources on the stuff we do. That is, our mental capacities won't diminish, we'll just redirect ourselves. I think that's probably likely to be the case most of the time, because while Hogg takes it too far, people in academia do value the experience of problem-solving in itself. They aren't likely to want to just stop doing that – indeed, some of them might even not be able to. But we also have to consider the temptations towards laziness, to jump straight to the final answer... and more insidiously, that maybe only the low-level stuff is sufficient to really keep the brain working at peak ability or prevent it from degrading. 

People on different sides of the AI debate all have very different views on this. I lean towards "this will just be a good thing, most of the time". My personal experience is this is something which really lets us get shit done, and nothing remotely comparable in getting-shit-done abilities has preceded this in my lifetime. Not even close. It's hard not to be optimistic about that, and I incline towards the view that a thing which is good for getting shit done is highly unlikely to actually decrease the amount of getting shit done. Even so, I don't think the effects are fully predictable, and I don't dismiss the tendency to skip the legwork* and get to the answer instead.

* Isn't skipping itself legwork though ?

I also have to recognise the special privilege of astronomy here. It's not quite that our results are of no importance, as Hogg claims. It's a question of precision and timescales. On the long term, our broad results are as important as anything else in any field of knowledge. But the exact values we determine in the short term are indeed of no importance to anyone else except ourselves. This frees us from any immediate need to be productive : our findings almost certainly won't cure cancer or relieve pain or solve famine, except possibly through spin-offs. So this means we can, and perhaps will, continue to do some of the low-level stuff genuinely for enjoyment, just as I've written this extremely long blog entirely by hand* because it's something I wanted to do, not because I think more than half-a-dozen people are likely to read it or even because I thought I could do a better job than a chatbot.

* I deliberately added a keyboard shortcut to make typing en dashes easier, so don't let those fool you. I happen to like en dashes, mmkay ?

The importance of small data in astronomy remains key to allowing humans to make genuine contributions, just as amateur astronomers can and do still make valuable discoveries for the professionals. This is not something that has direct equivalences in other fields, but it does give us some clues to the (short-term) future. While Hogg may be right that an LLM can write a paper 100,000 times faster than a human, they don't have any innate desire to do so. There's no topic an LLM actually has an interest in because they literally don't exist until prompted by a human : they have a crude agency, but no consciousness. And yes, while it might eventually be possible the deploy the large-scale compute Hogg hypothesises* could lead to factors more in the billions, where we could simply ask, "Please solve all problems in astronomy" and get something back that wasn't drivel, I will go so far as to say this isn't happening this decade.

* If I were him, I would make it my professional mission to have a hypothesis named after me.

So we're safe for the foreseeable future. The techbro predictions are hype, but they're not made up of nothing. AI is and will be transformative for astronomy. Anyone thinking that we can ignore it, that things will carry on as normal, or even that they can clearly see where this is going, well, enjoy your blissful ignorance, but I'm afraid you didn't get the memo. My only prediction is that the solution will be obvious after the fact and a handful of people who got lucky with the right call will proclaim themselves wise sages... unless they too are replaced with killer robots. Only time will tell.

Tuesday, 16 June 2026

AI Can Help Us Publish Less, Says Scientist

Not really all that much about AI in this one, actually.

I think everyone agrees that "publish or perish" is bad, but I don't think the approach suggested here makes much sense. However, I do like the following point very much :

AI is entering a publication system already swollen by proliferation, marked by signs of declining disruptiveness, and under growing pressure at the level of review and evaluation. Under those conditions, scientific papers can acquire negative epistemic value: not because they are wrong, but because the understanding they add no longer compensates for the time and attention they draw away from editors, referees, readers and colleagues trying to place them within what is already known.

Once papers and citations become central to hiring, promotion and funding, while publication itself also becomes a commercial object,  proliferation acquires a force of its own. Science fills with large bubbles of urgent-but-not-important writing: work that is timely, legible to evaluators, easy to package and profitable to circulate. 

When scarce time and judgment are drained by papers whose contribution no longer justifies what they demand from everyone else, they contribute negatively to the collective production of knowledge, slowing and hampering the rise of the mountain.

Well, I agree ! The author is also careful to say that having lots of incremental papers is not in itself a bad thing. But we're reaching the point where trying to maintain the vast wealth of relevant knowledge needed for a small amount of progress is outlandish. We require typically 15-20 pages or more to describe in meticulous detail what was done, why it was done etc. etc. etc. all for the sake of a small advancement that nobody will care about except a handful or direct competitors who will tear it apart limb from limb... we're burning the candle at both ends, making an enormous amount of work for ourselves both when reading and publishing.

I exaggerate, but only slightly. We also have to spend an inordinate amount of time adhering to strict and and utterly pointless journal standards which do exactly nothing to advance the state of the field and do an awful lot towards making the final paper less readable.

So I agree with the diagnosis. I'm much more skeptical of the suggested treatment, not because I have anything against AI in science (quite the opposite !), but because I don't think this is right approach and won't do anything much to address the problem.

The first concerns the visibility of non-paper contributions, such as code on GitHub, data on Zenodo, curated benchmarks, public notebooks, reproducibility packages, and living syntheses. They have existed for years and are valued by practising scientists. But they have remained second-class citizens: largely invisible to hiring and evaluation committees, and usually legitimized only via a paper that describes them rather than recognized in their own right. AI can change that communication layer. 

Can it though ? Only very weakly, I think. It can make a code or other product more intelligible to evaluators, but this won't mean anything if they don't have some box to tick on their reports.

The second concerns time. AI can absorb much of the routine labour that now consumes researchers’ effort: literature mapping, code scaffolding, documentation, reproducibility checks, exploratory work on alternative paths, first-pass synthesis across neighbouring literatures... It is that it can remove some of the weight that now pushes many worthwhile directions out of reach.

I agree with that one. AI – and I mean AI from the last few months or so with very low hallucination rates – is very, very good at turning routine but unintelligible research into something accessible and useable by experts who aren't specialists in that particular field. It can also be trusted at least with grunt-work that produces easily-testable results. Don't want to spend time writing a GUI for your enormously complex script ? I completely get that, it's boring. But you can easily have an AI slap something on which, if not perfect, is still massively better at not having one at all, and can generally be refined to something of a decent standard fairly easily. Thus you end up producing stuff which is not only powerful, but actually useable and accessible to a wider audience instead of the hardcore loonies who insist that everything should be done via the command line for some reason.

The third concerns evaluation. AI can strengthen the front end of review by helping editors, evaluation panels, and funders with triage, novelty checks, literature comparison, technical consistency checks and the detection of obvious pitfalls... This is one of the places where AI can directly counter the danger of negative epistemic value. When scarce time is spent processing papers that add too little in return, knowledge suffers. Review would still take time where needed. What could change is the amount of low-level labour surrounding it, so that more of the community’s limited attention is reserved for contributions that genuinely deserve it.

Here I think this is basically true, but not so much for review itself as for distilling knowledge into what researcher's actually need. I still want a human expert reviewing the paper and checking the whole thing carefully for errors, but once published, AI is very powerful for checking on whether a paper actually contains something I actually need to use. It's just not reasonable to expect authors to fully read hundreds of 20+ page papers in full when producing their own; the vast majority of citations are selected only because of one or two key results in each paper, not because everyone is reading absolutely everything.


I've said it before and I'll say it again. What we need are two main changes, one at the level of the journals, and the second at the level of evaluation. The two are inextricably linked. Instead of publishing just in a regular journal or Science/Nature (i.e. ordinary versus prestigious), we need far more journals and divisions within journals. We need to actively demark papers that required a shittonne of work from those which were rattled off in an afternoon. There is real value in disseminating pure ideas with absolutely no testing, but such a paper shouldn't be held to the same standard as the results of running a huge simulation or cataloguing an enormous set of observations. And we need, therefore, to insist that these different levels of papers – which need different standards of review rigour, clearly and publicly stated (most journals do nothing of the sort, never stating what the reviewer is actually supposed to do and when they should shut up) – are actually accounted for in evaluations. Maybe your department already has lots of hard-working incrementalists and needs someone more creative. Maybe it's the opposite. All have value in the right context

(Incidentally, my institute does account for non paper-producing duties in our internal evaluations, but I've yet to see this much used in external applications like jobs and grants)

A closely related point is that we probably need to think more about how we want papers to be structured. The prestigious journals tend toward a much more readable format : here are the key results together with the primary reasoning and potential pitfalls in this 6-page report, and here, in this 20+ page appendix, are the deep technical details of how we did it. This makes it massively easier to read the key results if you don't need the gory details, and has no real downsides if the technical stuff is what you're after.

So : papers which are easier to read; papers reviewed to different standards and with different labelled metrics; and evaluations which account for the different value that different types of product bring to the table. That would help a great deal, I think : not so much by publishing less as publishing differently and recasting what it is we actually have to read. AI might help here, but only as a second-order effect. It's not the main route, in my view, to beating "publish or perish" culture to its deserved death with a big stick.

Wednesday, 18 March 2026

The Secret History of Dark Matter

One of the nice things about doing all my academic reading on a digital tablet is that I can download papers I'd like to read purely for the sake of interest. I don't get much chance to actually read them, but it's better than having a huge list of bookmarks I'll never check, or a stack of printed papers so large it could qualify as a carbon sequestering facility.

Finally, I managed to get round to reading one such paper, and it turned out to be a thoroughly worthwhile read. Maybe a bit on the lengthy side, but then that's what this blog is for.

Anyway, the popular narrative history of dark matter goes something like this. Jan Oort and Fritz Zwicky made some early claims that they might have found it back in the 1930s, but it was all from purely weak observational evidence. There wasn't any particular theoretical reason for it, so everyone ignored it until the 1970s when Vera Rubin and others started finding that galaxies were rotating much too quickly. Et voila, paradigm shift, everyone got very excited, and this resulted in the modern cosmology we know and love.

This paper* makes some important revisions to what actually happened. Rather than being pure observational luck, Zwicky had clear theoretical motivation for dark matter. Some of his underlying reasons for expecting dark matter have long since been thoroughly refuted, but some have intriguing similar aspects to modern theories. And he was even, quite likely, actively searching for it, with his technique being one that's still in regular use today. His sample size was absolutely shite – seven or eight galaxies would never be enough to convince anyone, but the method was sound. 

* I'm unsure of the provenance of the article. It appears to be only uploaded to preprint servers with no hint of whether it's submitted to a journal or not. 

What he did not do was stumble on a result he couldn't explain and invent the idea of dark matter as an ad hoc "fudge factor". More on this later, but even from its earliest days, there were already strong theoretical reasons to believe dark matter existed before observations started to get ahead of the game. Today, everyone knows about Zwicky, but most people forget the theoretical paradigms in which he operated. This paper attempts to set the record straight.

What follows is my summary of the paper. I've tried only to shorten and simplify the content rather than put too much of my own spin on things.




Zwicky did not at all like the idea of an expanding Universe. Today, this is as well-established as any result in science can be, but at the time, he had good reasons to be skeptical. The difference in redshifts caused by expansion of space and motion through space were not yet fully appreciated, and the "breathtaking speeds" of galaxies moving at thousands of kilometres per second therefore seemed ludicrous. From a contemporary vantage point this seems weird, but when you only have a meagre handful of data points and they seem to be implying that something outlandish is happening, most of the time it's actually quite sensible to bet on your pre-existing ideas. 

This is what led to Zwicky's "tired light" hypothesis as an alternative to cosmological expansion. The idea was that photons would lose energy as they travelled for long distances, becoming redder and redder. Galaxies might indeed be at stupendous distances – I see no indication Zwicky ever doubted this monumental discovery from Hubble – but their speeds might be an illusion. 

The thing is, what would cause a photon to become tired ? Zwicky's answer was that there must be some intervening material, unseen through direct observation, but inferable through its effects on photons. And he wasn't the first to suggest that there could be some quantity of dark matter out there, with the authors suggesting that actually the opposite hypothesis – the idea that all matter must be luminous – was regarded as equally audacious.

The prehistory of dark matter is long and complicated. And it really is prehistory, because it really seems to be only much later that we get to the idea of a genuinely new type of substance, the concept of a material that only interacts with our own through gravity. None of the earliest ideas ever suggested it was anything other than matter which was perfectly normal, just in a state where it was bloody difficult to see. This includes things like Mitchell's "dark stars", which generated photons but which were trapped by the star's massive gravitational field; unseen planets like Neptune that just hadn't been spotted yet; most crucially of all, Einstein and de Sitter's cosmology required dark matter to maintain a flat Universe.

So Zwicky's motivation for dark matter seems to be neither quite that he was the curmudgeonly contrarian of popular lore (though he definitely was a cantankerous git) nor a visionary ahead of his time. Rather he was operating in conditions which are not really directly comparable to the modern scientific view at all. He had three main motivations for expecting dark matter, and these are best understood on their own terms. First, he thought this could explain away the expanding universe, which he viewed as a problem rather than a reality*, through his tired light hypothesis. Second, he knew dark matter could greatly help with the Einstein de Sitter (EdS) model, which Zwicky seems to have favoured. And third, he also thought it could explain the origin of cosmic rays.

* This is much my own view on dark matter. I don't quite understand why many people treat it as apparently obviously problematic and in need for explanation, rather than accepting that this is just what the data shows. I'm simplifying here, but you get the point.

Tired light is well known, and long since refuted, but the problem of cosmic rays is more often forgotten. Nobody understood where these high energy particles were coming from, but since they appeared to be uniform across the sky, they either had to be from something very nearby or very far away indeed. Since there were no obvious nearby sources (that is, within our own small patch of the Milky Way) that might explain them, a cosmic origin was very reasonable. Not knowing about active galactic nuclei and the like, Zwicky's dark matter seemed like a pretty good site for their origin. I have to say I don't quite understand the author's argument as to why Zwicky didn't believe they could originate from anything luminous, but still, if they originated from low density matter filling all of space, this would naturally explain their uniform distribution across the sky.

He also had a good reason to search in clusters. Lord Kelvin had suggested in 1904 that galaxy dynamics could give clues to their total mass independently of their brightness, and Zwicky extended this to clusters (incidentally he also realised that, in partial contradiction to Einstein, a galaxy cluster might be so massive that gravitational lensing there might be so strong as to be detectable).

This meant that Zwicky had very solid grounds for targeting Coma. He doesn't set out his motivations in his own paper, so all this is a contextual reading by the modern authors, but in my view it's a compelling one. Zwicky had good reasons to believe dark matter could be detected in clusters by examining their dynamics, and similarly convincing arguments for its existence which were completely independent of dynamical considerations. It's entirely credible that his observations were done deliberately for this very purpose, and his putative discovery of dark matter wasn't something that he just happened to have stumbled on by chance at all. 

The most interesting aspect of this to me is that dark matter was a prediction of relativity. I've long wondered if an early detection of dark matter would have throttled relativity in the cradle, but the answer from this paper is a clear "no"... but not for reasons we could still justify today. The thing about relativity is that it required dark matter for a flat Universe, but the amount required would have been far larger than in our modern estimates. Today's cosmology uses dark matter to explain the dynamics of galaxies and clusters, but the cosmological constant (and inflation) to keep the universe flat. At the time, it seems the equivalent mass density of the constant wasn't considered to be sufficient to do the job of flattening space. 

And it's important to remember that Zwicky was still making one hell of an extrapolation. From his seven or eight galaxies, he assumed that the rest of the Coma cluster galaxies (hundreds strong) were in stable equilibrium, and thus derive a value for the total mass which happened to be in agreement with the predicted dark matter content needed for the EdS model. But given that even Hubble's constant was not at all well-constrained at the time (Hubble's first value as 500 km/s/Mpc; today we think it's around 71 km/s/Mpc), the error bars on this were massive. So again, relativity predicted dark matter, but not directly, and not for reasons we can now sustain. History turns out to have been more complicated than my counterfactual musings.




The final part of the paper is more philosophical, considering whether Zwicky's idea really constitutes an ad hoc hypothesis. Certainly it seems not to have been something he just invented on the fly; he may have even been deliberately searching for it. As far as Zwicky's particular idea goes, the answer here is a decisive "no". He may very well have been doing the classic scientific model of hypothesis testing, with his observations set in a clear theoretical framework – and was definitely not trying to save Newtonian gravity from relativity, as has been claimed.

What about ad hoc hypotheses more general ? The paper gives quite a thorough discussion on the different perspectives on these, noting that the existence of Neptune was arguably just such a case. Explaining the orbits of the other planets was difficult without an extra one hitherto undetected, but there was no other good reason to expect the existence of such an object. And of course Neptune did turn out to exist, which completely scuppers the notion that "ad hoc" automatically means "wrong". 

In some extreme views, there are no ad hoc hypotheses at all : they can't be clearly defined and depend too much on circumstance, and since they can turn out to be right, there's no point in distinguishing them from any other hypothesis. The authors here note that both proponents of modern dark matter and those of modified gravity view the other as embracing ad hoc hypotheses in a pejorative sense : to dark matter adherents, modified gravity does nothing except explain rotation curves; to modified gravity researchers, dark matter does nothing except... explain rotation curves. And such hypotheses can be both conservative (seeking to save existing ideas, like Neptune in a Newtonian framework) and progressive (like Zwicky's dark matter in an EdS universe).

I think my take remains that the most important thing for a hypotheses is testability rather than how many ideas it explains. True, we might get a bit suspicious if an idea is invoked to explain a single, unique observable, but this is all we should do, rather than insisting the idea was no good. If you have no other grounds to suspect something, invent it on the fly to explain just one thing, and don't have any reason to expect you'll be able to use your idea elsewhere... then your idea might just be of low overall importance rather than actually wrong. It may, in fact, be perfectly reasonable to suspect the existence of a particular planet or galaxy based on observational evidence, and this may be of locally extreme importance : it just isn't likely to alter anything fundamental.

Where the concept of an ad hoc hypothesis does start to become more problematic, I think, is where it is invoked to explain the fundamental basis of a theory. If you need it to explain a single observation, but without this the whole theory collapses, then this should give you pause for thought. By no means does it suggest the hypothesis is wrong, but it's clearly better if your idea explains multiple things or a general situation rather than just one specific thing.

Neither the modern or Zwicky's concept of dark matter constitute such a thing. Both were and are used to explain multiple lines of evidence. In Zwicky's case, some of those aspects were his own ideas but some were completely independent. In that sense, say the authors, Zwicky should be recognised as the "quantifier, not the discoverer, of dark matter". He used methods both of his own and others devising to explain both his own and others observations ; the idea of dark matter itself was not original to Zwicky. Here irony piles atop irony. His findings were correct but unconvincing, with his paper not cited for 25 years; his value agreed with a theoretical framework which turned out to be completely wrong; his basic technique correct but involving a wild extrapolation. 

In this reading, Zwicky comes through as both a revolutionary and a staunch conservative. He had the best of ideas, he had the worse of ideas. But, pretty decisively, it seems we should give up on any misconception he simply invented dark matter to explain a few errant galaxies. Rather, in this particular case, he seemed to have been doing good science, the best he could do at the time – subsequently revised, but that's exactly what any good scientist can hope for. Zwicky did indeed push the boundaries forward, and if others had paid a bit more attention, the history of cosmology could have been completely different.

Thursday, 28 August 2025

Signal boost

One man's trash is another man's treasure, or so the old saying goes. And in astronomy, one man's signal is another man's noise. 

The classic example of this is dust in the Milky Way. If you're interested in dust, then you're a deeply weird person... or just interested in star formation. Dust "grains", which are actually about the size of smoke particles, are thought to be critical sites for star formation because they allow atomic gas to lose energy, cool, and collide with each other to form molecular gas. This is much denser than atomic gas, so eventually this can lead to the cloud collapsing to form a star.

But if you're less of a weirdo, dust just gets in the way. It ruins our majestic sky by blocking our view of all the stars, especially along the plane of the disc of the Galaxy. Some regions are much worse than others, but it's present at some level pretty much everywhere across the sky.

In radio astronomy we have a much more subtle and interesting problem. Of course we always want our observations to be as deep and sensitive as possible. But sometimes, it turns out, the noise in our data can actually be to our advantage – though it comes at a price.

Consider a typical spectrum of a galaxy as detected in the HI line. If you aren't familiar with this, take a look at my webpage if you want details. Basically, it shows us how bright the gas from a galaxy is emitting (that is, how dense it is) at any particular velocity. Even without knowing this rudimentary bit of information though, you can probably immediately identify the feature of interest in the signal :

All the spectra shown in this post are artificial, generated with a simple online code you can use yourself here.

We don't need to worry here about why the signal from the galaxy has the particular structure that it does. No, what I want to talk about today is the noise. That's those random variations outside the big bright bit in the middle.

This example shows a pretty nice detection. It's easy to see exactly where the profile of the galaxy ends and the noise begins. But even within the galaxy, you can see those variations are still present : they're just lifted up to higher values by the flux from the galaxy. Basically the galaxy's signal is simply added to the noise.

Now if you still have an analogue radio, you'll know that if you don't get the tuning just right, you'll hear the sounds from your station but only against a loud and annoying background hiss. The worse the tuning, the worse the noise. So you might well think that the following claim is more than a little dubious :

Fainter signals can be easier to detect in noisier data.

So counterintuitive is this that one referee said it "makes no sense whatsoever", doubling down to label it "bizarre" and "not just counterintuitive, but nonsensical".

This is wrong. I'll point out that the claim comes not from me but from Virginia Kilborn's PhD thesis  (now a senior professor). So how does it work ?

The answer is actually very simple : signal is added to noise. That is, regardless of how noisy the data is, the signal from the galaxy is still there. Let me try and do this one illustratively. Suppose we have a pure signal, completely devoid of noise, and for argument's stake we'll give it a top-hat profile (about a quarter of galaxies have this shape, so this isn't anything unusual) :

The "S/N" axis measures the signal to noise, a measure of how bright things look given the sensitivity of the data. The numbers in this case are garbage because I set the noise to zero.

Now let's add it to two different sets of noise, purely random (Gaussian), of exactly the same statistical strength but just different in their exact channel-to-channel values :


Note that even with this purely random noise, you can still see apparent ripples and variations in the baseline outside the source : even random noise, to the human eye, looks structured.


Oh ! What happened there ? Why is the second signal so much clearer than the first ? You can still see the first one, to be sure, but it's marginal, and could easily be mistaken for some weird structure in the baseline. The second isn't great either, but it looks a lot better than the first.

Noise is typically random. That means that some parts of it will have bits of higher flux while other parts will be a bit lower. If we add our signal to the higher flux bits, the total apparent flux in our source gets higher. That is, the real flux in our source obviously doesn't change, but what we would measure would be greater than if the noise wasn't there. And of course the opposite could happen too : we could have noise dimming that makes the source harder to detect as well as easier.

Noise boosting (shown in the carefully-chosen example above), on the other hand, is no less important but far less expected. Every once in a while, a faint signal will happen to align with some bright parts of the noise and turn a marginal signal into a clearer one. This doesn't really work in the sorts of audio radio signals you get on a household radio set as these are much too complex, but the signals of a galaxy are a good deal simpler. And all we need to detect them is (for this basic example at least) pure flux, which the noise can readily provide.

(As an aside, you might notice that the the actual peak levels in these cases aren't much different, though the average level inside the source profile is higher in the second case. While peak levels most certainly can be affected by noise boosting and dimming, what's absolutely crucial here is what detection method we're using to find the signals. I'll return to this below.)

Of course, there are limits to this : it will only work for signals which are comparable in strength to the noise. As the noise level gets higher, the random variations will increasingly tend to "wash out" our signal. Now the level of the signals we can receive varies hugely depending on the nature of our data set, but the signal-to-noise ratios (S/N or sometimes SNR) tend to be fairly constant. That is, a signal which is ten times the typical noise value (the rms) has the same statistical significance in any data set, but the actual flux value it corresponds to can be totally different. But expressing signal strength in terms of the noise level makes things very convenient : a five sigma (5σ) source just means something that's five times brighter than the typical noise level.

So suppose we have a 2σ source which happens to align with a 3σ peak in the noise : bam, we've got ourselves a quite respectable 5σ detection*. But if we keep the flux level of the signal we're adding the same and increase the noise level, then the ratio of S/N will go down. Instead of adding 2σ to 3σ, we'll be adding ever lower and lower values : the "sigma level" of the noise won't change, but that of the signal certainly will. Pretty quickly we won't be shifting that 3σ peak from the noise by any appreciable degree. We'll be adding the same flux value but to an ever-greater starting level. 

* Sometimes five sigma is quoted as a sort of scientifically universal gold-standard discovery threshold. This is simply not true at all, because if you have enough data, you'll get that level of signal just by chance alone. Far more importantly, the noise in real data is often far from being purely random, so choosing a robust discovery threshold requires a good knowledge of the characteristics of the data set.

As a possibly pointless analogy, consider lions. If you go from having no lions to one lion, you've just put yourself in infinitely more danger. If you add a second lion you're in even more trouble. But if you've got ten lions and add one more, you won't really notice the difference.

"But hang on," you might say, "surely that means that your earlier claim that fainter signals are more detectable in nosier data can't possibly be correct ?". A perfectly valid question ! The answer is that it depends on how we go about detecting the signals. The details are endless, but two basic techniques are to search for either a S/N ratio or a simple flux threshold. These can give very different results to each other.

Now if you use S/N, which is generally a good idea, then indeed signals of lower flux levels generally don't do well in increasingly nosier data because of the ever-smaller relative increase in the signal. And of course, the effect of the noise to not merely obscure but actually suppress the signal will get ever greater, since there's just as much chance of aligning with a low-value region of the noise as a high-value region. 

But S/N is not the only way of detecting signals. You might opt instead to use a simple flux threshold instead : it's computationally cheaper, easier to program, and most importantly of all it gives you more physically meaningful results. If you do it that way, then it's a different story. When you add a signal to noise the flux level always increases, making it much easier to push the flux above your detection threshold by this method. Which makes noise boosting very much easier to explain.

Note here the change of axes values compared to the previous examples. All I did was increase the noise level by ~25% and here both peak flux and S/N levels have increased. It might not look easier to detect visually, but statistically, by some measures this one is more significant than the previous cases !

Even more interesting is the so-called Eddington bias. What this means is that any survey will tend to overestimate the flux of its weakest signals : those signals which are so faint that they can only be detected at all thanks to chance alignments with the noise. In this case, that when someone comes along and does a deeper survey, they'll often find that those sources have less flux than reported in the earlier, less-sensitive data.

There are plenty of other subtleties. The signal might not need to be perfectly superimposed on a noise peak : if it's merely adjacent to it, that can create the appearance of a wider, brighter signal which can be easier to detect to some algorithms (and people !). And of course while we'd like noise to be perfectly uniform and random, this isn't always the case. Importantly, the rms value doesn't tell us anything at all about the coherency of structures in the data, as so powerfully shown by the ferocious Datasaurus

For the eye the effects of this can be extremely complex, and are poorly understood in astronomy. If you have very few coherent noise structures, for example, you might think this nice clean background would make fainter structures easier to spot. But actually, I have some tentative evidence that this isn't always the case : the eye can be lulled into a false sense of emptiness, whereas if there are a few obvious structures to attract attention, you start to believe that things are present so you're more likely to identify structures. No doubt if your data was dominated by structures then the eye would, in effect, perceive them as background noise again and the effect would diminish, but this is something that needs more investigation. My guess is there's a zone in which you have enough to encourage a search but few enough that they don't obscure the view. 

Ultimately, getting deeper data is always the better option. For every source noise-boosted to detectability, there'll be another which is suppressed and hidden. But this explains very neatly that supposedly "nonsensical" result*, especially if your source-finding routine is based on peak flux : of course if you add signal to noise and have the same flux threshold in your search, you're more likely to find the fainter signals in the nosier data... up to a point. Set your threshold too low and you'll just find spurious detections galore, but hit the sweet spot and you'll find signals you otherwise couldn't.

* Kudos to the referee for accepting the explanation; I never heard of any of this until a couple of years ago either. Extragalactic astronomy is full of stuff which isn't that difficult but seems feckin' confusing when you first encounter it because it isn't formally taught in any lectures !

There's nothing weird about noise boosting then. Mathematically it makes complete sense. But when you first hear about it it sounds perplexing, which just goes to show how deceptively simple radio astronomy can be. Noise boosting, at least at a basic level, is quite simple, but simple isn't the same as intuitive.

Wednesday, 19 June 2024

The shoe's on the other foot

My, how the tables have turned. The hunter has become the hunted. And various other cliché's indicating that the normal state of affairs have become reversed.

That is, as well as having to write an observing proposal, I find myself for the first time having to review them. Oh, I've reviewed papers before, but never observing proposals. This came about because ALMA has a distributed proposal review system : everyone who submits their own proposal has to review ten others. And since this year I finally submitted one, I get to experience this process first hand.


The ALMA DPR procedure

When you submit a proposal, you indicate your areas of expertise and any conflicts of interests  – collaborators and direct competitors who shouldn't be reviewing your proposal, either because they'd stand to benefit from it being accepted or would love to take you down a peg. It's a double-blind procedure : your proposal can't contain any identifying information and you don't know who the reviewers are. Some automatic checks are also carried out to prevent Co-Is on recent ALMA proposals being assigned as reviewers, and suchlike.

Then your proposal is sent off for initial checks and distributed to ten other would-be observers who also submitted observing proposals in the current cycle. You, in turn, get ten proposals to review yourself. Each document is four pages of science justification (of which normally one or even two pages are taken up with figures, references, and tables) plus an unlimited-length technical section containing the observing parameters for each source plus some brief justification on the specifics (in practise, in most proposals each of these so-called "science goals" are very similar, using the same observing setup on multiple targets). You then write a short review of each one, of a maximum of 4,000 characters but typically more like ~1,000 (or even less) describing both the strengths and weaknesses of each. You also rank them all relative to each other, from 1 (the strongest) to 10 (the weakest).

That's stage one. A few weeks later, in stage two you get to see everyone else's reviews for the same proposals, and can then change your own reviews and/or rankings accordingly, if you want to. So far as I know, each reviewer gets a unique group of ten proposals to review, so no two reviewers review the same set of proposals, meaning you can't see the others rankings. Exactly how their rankings are then all compared and combined, and ultimately, translated into awarded telescope time, remains a mystery to me. Those details I leave for some other time, and I won't go into the details of anonymity* here either : I seem to recall hearing that this gives a better balance of both experience and gender, but I don't have anything to hand to back this up.

* I will of course continue to respect the anonymity requirements here, and not give any information that could possibly identify me as anyone's reviewer.

Instead I want to give some more general reflections on the process. To be honest I went into this feeling rather biased, having received too many referee comments which were just objectively bollocks. I was quite prepared to believe the whole thing would be essentially little better than random, which is not a position without merit.


First thoughts

And my initial impressions justified this. It seemed clear to me that everyone had chosen interesting targets and would definitely be able to get something interesting our of their data, making this review process a complete waste of time.

But after I let things sink in a bit more, after I read the proposals a bit more carefully and made some notes, I realised this wasn't really the case. I still stand by (with one exception) that all proposals would result in good science, but the more I thought about it, the more I came to the conclusion that I could make a meaningful judgement on each one. I tried not to judge too much whether one would do better science than another, because who am I to say what's better science ? Why should I determine if studies of extrasolar planets are more important than active galactic nuclei ?

These aren't real examples, but you get the idea. Actually the proposals were all aligned very much more closely with my area of expertise. The length of four pages I would say is "about right", it gave enough background for me to set each proposal in its proper context as well as going into the specific objectives.

Instead, what I tried to assess was whether each project would actually be able to accomplish the science it was trying to do. I looked at how impactful this would be only as a secondary consideration. There isn't really any right or wrong answer as to whether it's better to look at a single unique target versus a statistical study of more typical objects, but I tried to judge how much impact the observations would likely have on the field, how much legacy value they would have for the community. But first and foremost, I considered whether I was persuaded the stated science objectives could actually be carried out if the observations themselves reached their design spec.


Judgement Day

And this I found was something I could definitely judge. Two proposals to me stood out as exemplary, perfectly stating exactly what they wanted to do and why, exactly what they'd be able to achieve with this. It was very clear that they understood the scientific background as well as anyone did. I initially ranked these essentially as a coin-toss as to who got first and who got second place; I couldn't meaningfully choose between them.

At the opposite extreme were two or three which didn't convince me at all. One of the principle objectives of one of them was just not feasible with the data they were trying to obtain, and they themselves presented better data in their proposal that they already had which would have been much more suitable for this. Lacking self-consistency is a black mark in just about any school of thought. Another looked like it would observe a perfectly good set of objects, but contained so many rudimentary scientific errors that there was no way I could believe they'd do what they said they would do. 

Again, deciding which one to rank lowest was essentially random, though I confess that one of them just wound me up the wrong way more than the other.

In the middle were a very mixed bunch indeed. Some had outstanding ideas for scientific discovery but were very badly-expressed, saying the same thing over and over again to the nth degree (I would say to these people, there's no obligation to use the full four pages, and we should stipulate this in the guidance to observers and reviewers alike. I tried to ignore the poor writing style of these and rank them highly because of the science). Some oversold the importance of what they'd do, making unwarranted extrapolations from their observations to much more general conclusions. Some had a basically good sample but claimed it was something which it clearly wasn't; others clearly stated what their sample was but the objects themselves were not properly representative of what they were trying to achieve.

This middle group... honestly here, a random lottery would work well. On the other hand, there doesn't seem any obvious reason not to use human judgement here either, because for me at least this felt like a random decision anyway. And if other people's judgements are similar then clearly there are non-random effects which probably should actually be accounted for, whereas if they are truly random then the effects will average out. So there's potentially a benefit in one case and no harm in the other, and in any case there almost certainly is a large degree of randomness at work anyway.


Reviewing the reviews

I went through a similar process of revising my expectations in stage 2, though to a lesser degree. At first glance I didn't think I'd need to change my reviews or rankings, but on carefully checking one of the other reviews, I realised this was not the case. One reviewer out of the ten had managed to spot a deeply problematic technical issue in one of the proposals that I otherwise would have ranked very highly. And on checking I was forced to conclude that they were correct and had to downgrade my ranking significantly. This alone makes the process worth doing : 1 out of 10 is not high, but with ~1,600 proposals in total, this is potentially a significant number overall.

Reading the other reviews turned out to be more interesting than I expected. While some did raise exactly the same issues with some of the proposals that I had mentioned, many didn't. Some said "no weaknesses" to proposals I thought were full of holes. One even said words to the effect that "no-one should doubt the observers will do good science with this", a statement I felt presumptuous, biased, and bordering on an argument-from-authority : it's for us the reviewers to decide this independently; being told what we should think is surely missing the point. 

The reverse of this is that some proposals I though were strong others thought were weak – very weak, in some cases. Everyone picks up on different things they think are important. There was one strange tendency for reviewers to point out that the ALMA data wouldn't be of the same quality as comparison data. This is fine, except that the ALMA data would usually have been of better quality, and downgrading it to the same standard is trivial ! I sort of wished I'd edited my reviews to point this out. Some also made comments on statistics and uncertainties that I thought were so generic as to be unfair, yes of course things might be different from expectations, but that's why we need to do observations !

What the DPR doesn't really do is give any chance for discussion. You can read the other reviews but you can't interact with the other reviewers. It might have been nice to have somewhere where we could enter a "comment to other reviewers", directed to the group, or at least have some form of alert system when reviews were altered. Being able to ask the observers questions might have been nice, but I do understand the need to keep things timely as well. On that front, reviews varied considerably in length; mine were on the longer and to be honest perhaps overly-long side (I think my longest was nearly 2,000 characters), while one was consistently and ludicrously short.

All this has given me very mixed feelings about my own proposal. On the one hand, I don't think it's anywhere near the worst, and I stand behind the scientific objectives. On the other, I think I concentrated overmuch on the science and not enough on the observational details. Ranking it myself with hindsight I'd probably have to put it in the lower third. It was always a long shot though, so I'll be neither surprised nor disappointed by the presumed rejection. One can but try with these things.

One thing I will applaud very strongly is the instruction to write both strengths and weaknesses of each proposal. All of them, bar none, had some really good points, but it was helpful to remind myself of this and not get carried away when reviewing the ones I didn't much like. Weaknesses were more of a mixed bag; one can always find something to criticise, although in some cases they aren't significant. Still, I found it very helpful to remember that this wasn't an exercise in pure fault-finding.




How in the world one judges which projects to actually undertake, though... that seems to me like the ultimate test of philosophy of science. Groups of experts of various levels have pronounced disagreements about factual statements; some notice entirely different things from others. There's the issue of not only will the science be significant, but also whether the data can be used in different ways from what's suggested. That to me remains the fundamental problem with the whole system, that one can nearly always expect some interesting results, but predicting what they could be is a fool's game.

Overall, I've found this a positive experience. Reading the full gamut of excellent to poor proposals really gives a clearer idea of what reviewers are looking for, something it's just not possible to get without direct experience. Not for the first time, I wonder a lot about Aumann's Agreement Theorem. If we the reviewers are rational, we ought to be persuaded by each other's arguments. But are we ? This at least could be assessed objectively, with detailed statistics possible on how many reviewers change their ranking when reading other reviews. 

And at the back of my mind is a constant paradoxical tension : a strong feeling that I'm right and others are wrong, coupled with the knowledge that other people are thinking the same thing about me. How do we reconcile this ? For my part, I simply can't. I formulate my judgement and let everyone else to the same, and hope to goodness the whole thing averages out to something that's approximately correct. The paradox is that this in no way makes me feel any the less convinced of my own judgments, even knowing that some fraction of them simply must be wrong.

Other aspects are much more tricky. This is a convergence of different efforts, both trying to asses what-is-true (what science claimed is factually correct, why do experienced experts still disagree on some points), what will likely benefit the community the most, and how we try and account for the inevitably uncertain and unpredictable findings. As I've said before many times, real, coal-face research is extremely messy. If it isn't already, then I would hope that telescope proposals ought to be an incredibly active field of research for philosophers of science.

Thursday, 6 June 2024

The data won't learn from itself

Today I want to briefly mention a couple of papers about AI in astronomy research. These tackle very different questions from the usual sort, which might examine how good LLMs can be at summarising documents or reading figures and the like. These, especially the second, are much more philosophical than that.

The first uses an LLM to construct a knowledge graph for astronomy, attempting to link different concepts together. The idea is to show how, at a very high level, astronomical thinking has shifted over time : what concepts were typically connected and how this has changed. Using distributional semantics, where the meanings of words in relation to other words are encoded as numerical vectors, they construct a very pretty diagram showing how different astronomical topics relate to each other. And it certainly does look very nice – you can even play with it online

It's quite fun to see how different concepts like galaxy and stellar physics relate to each other, how connected they are and how closely (or at least it would be if the damn thing would load faster). It's also interesting to see how different techniques have become more widely-used over time, with machine learning having soared in popularity in the last ten years. But exactly what the point of this is I'm not sure. It's nice to be able to visualise these things for the sake of aesthetics, but does this offer anything truly new ? I get the feeling it's like Hubble's Tuning Fork : nice to show, but nobody actually does anything with it because the graphical version doesn't offer anything that couldn't be conveyed with text.

Perhaps I'm wrong. I'd be more interested to see if such an approach could indicate which fields have benefited from methods that other fields aren't currently using, or more generally, to highlight possible multi-disciplinary approaches that have been thus far overlooked.


The second paper is far more provocative and interesting. It asks, quite bluntly, whether machine learning is a good thing for the natural sciences : this is very general, though astronomy seems to be the main focus. 

They begin by noting that machine learning is good for performance, not understanding. I agree, but once we do understand, then surely performance improvements are what we're after. Machine learning is good for quantification, not qualitative understanding and certainly not for proposing new concepts (LLMs might, and I stress might, be able to help with this). But it's a rather strange thing to examine, and possibly a bit of a straw man, since I've never heard of anyone thinking that ML could do this. And they admit that ML can be obviously beneficial in certain kinds of numerical problems, but this is still a bit strange : what, if any, qualitative problems is ML supposed to ever help with ?

Not that quantitative and qualitative are entirely separable. Sometimes once you obtain a number you can robustly exclude or confirm a particular model, so in that sense the qualitative requires the quantitative. But, as they rightly point out, as I have myself many times, interpretation is a human thing : machines know numbers but nothing else. More interestingly they note :  

The things we care about are almost never directly observable... In physics, for example, not only do the data exist, but so do forces, energies, momenta, charges, spacetime, wave functions, virtual particles, and much more. These entities are judged to exist in part because they are involved in the latent structure of the successful theories; almost none of them are direct observables. 

Well, this is something I've explored a lot on Decoherency (just go there and search for "triangles"). But I have to ask, what is the difference between an observation and a measurement ? For example we can see the effects of electrical charge by measuring, say, the deflection of a hair in the static field of a balloon, but we don't observe charge directly. But we also don't observe radio waves directly, yet we don't think they're less real than optical photons, which we do. Likewise some animals do appear to be able to sense charge and magnetic fields directly. In what sense, then, are these "real" and what sense are they just convenient labels we apply ?

I don't know. The extreme answer is that all we have are perceptions, i.e. labels, and no access to anything "real" at all, but this remains (in some ways) deeply unsatisfactory; again, see innumerable Decoherency posts on this, search for "neutral monism". Perhaps here it doesn't matter so much though. The point is that ML cannot extract any sort of qualitative parameters at all, whereas to humans these matter very much – regardless of their "realness" or otherwise. If you only quantify and never qualify, you aren't doing science, you're just constructing a mathematical model of the world : ultimately you might be able to interpolate perfectly but you'd have no extrapalatory power at all.

Tying in with this and perhaps less controversially are their statements regarding why some models are preferred over others :

When the expansion of the Universe was discovered, the discovery was important, but not because it permitted us to predict the values of the redshifts of new galaxies (though it did indeed permit that). The discovery was important because it told us previously unknown things about the age and evolution of the Universe, and it confirmed a prediction of general relativity, which is a theory of the latent structure of space and time. The discovery would not have been seen as important if Hubble and  Humason had instead announced that they had trained a deep multilayer perceptron that could predict the Doppler shifts of held-out extragalactic nebulae.

Yes ! Hubble needed the numbers to formulate an interpretation, but the numbers themselves don't interpret anything. A device or mathematical model capable of predicting the redshifts from other data, without saying why the redshifts take the values that they do, without relating it to any other physical quantities at all, would be mathematical magic, and not really science.

For another example, consider the discovery that the paths of the planets are ellipses, with the Sun at one focus. This discovery led to extremely precise predictions for data. It was critical to this discovery that the data be well explained by the theory. But that was not the primary consideration that made the nascent scientific community prefer the Keplerian model. After all, the Ptolemaic model preceding Kepler made equally accurate predictions of held-out data. Kepler’s model was preferred because it fit in with other ideas being developed at the same time, most notably heliocentrism.

A theory or explanation has to do much more than just explain the data in order to be widely accepted as true. In physics for example, a model — which, as we note, is almost always a model of latent structure — is judged to be good or strongly confirmed not only if it explains observed data. It ought to explain data in multiple domains, and it must connect in natural ways to other theories or principles (such as conservation laws and invariances) that are strongly confirmed themselves.  

General relativity was widely accepted by the community not primarily because it explained anomalous data (although it did explain some); it was adopted because, in addition to explaining (a tiny bit of new) data, it also had good structure, it resolved conceptual paradoxes in the pre-existing theory of gravity, and it was consistent with emerging ideas of field theory and geometry.

Which is a nice summary. Some time ago I'd almost finished a draft of a much longer post based on this this far more detailed paper which considers the same issues, but then blogger lost it all and I haven't gotten around to re-writing the bloody thing. I may yet try. Anyway the need for self-consistency is important, and doesn't throttle new theories in their infancy as you might expect : there are ways to overturn established findings independent of the models. 

The rest of the paper is more-or-less in line with my initial expectations. ML is great, they say, when only quantification is needed : when a correlation is interesting regardless of causation, or when you want to find outliers. So long as the causative factors are well-understood (and sometimes they are !) it can be a powerful tool for rapidly finding trends in the data and points which don't match the rest. 

If the trends are not well-understood ahead of time, it can reinforce biases, in particular confirmation bias by matching what was expected in advance. Similarly, if there are rival explanations possible, ML doesn't help you choose between them if they don't predict anything significantly different. But often, no understanding is necessary. To remove the background variations in a telescope's image it isn't necessary even to know where all the variations come from : it's usually obvious that they are artifacts, and all you need to is the mathematical description of them. Or more colourfully, "You do not have to understand your customers to make plenty of revenue off of them." 

Wise words. Less wise, perhaps only intended as a joke, are the comments about "the unreasonable effectiveness of ML", that it's remarkable that these industrial-grade mathematical processes are any good for situations to which they were never designed. But I never even got around to blogging Wigner's famous "unreasonable effectiveness" essay because it seemed worryingly silly. 

Finally, they note that it might be better if natural sciences were to shift their focus away from theories and more towards the data, and that the degeneracies in the sciences undermine the "realism" of the models. Well, you do you : it's provocative, but on this occasion, I shall allow myself not to be provoked. Shut up and calculate ? Nah. Shut up and contemplate.

Why Bother ?

It's rare that I manage to read any longer pieces on arXiv that aren't strictly about galaxy evolution, but today I indulge myself. ...