Smut Clyde continues exploring things which nobody else would ever want to explore, reading papers which nobody else would ever want to read. You are now to face more of nonsense machine learning based on nonsense template made from stolen and garbled datasets, but fear not – it all passed peer review!
What follows is a follow-up to Smut’s apparently insufficiently disturbing previous story about those AI training datasets on Kaggle.

The subtleties a spectrograph would miss
by Smut Clyde
The ‘Kaggle Dance’ is a form of terpsichorean signalling used by workers in a papermill to pass along to their hive-mates the direction of the latest data-repository they found, along with information (encoded by shaking their abdomens) about its richness and potential for fake-paper exploitation…
No, wait, that’s the Waggle Dance. Still, it serves to warn readers that this will be a follow-up to an earlier post about papermills strip-mining the resources of Open Science, for of course there were out-takes!
But first I must pass on the recent discovery that frequent PubPeer contributor Hoya camphorifolia is an AI. Full credit to Mohit Bajaj and his colleague Arvind Singh for this realisation. Presumably H. camphorifolia has a staff of humans who can click on the “I am not a robot” bot-blocking button on its behalf when it needs to access a paper.
“it seems the AI used for an anonymous commenter is not able to read the article and randomly comments on the article. Also, some mistakes are bound to happen, which cannot be considered for some anonymous person to put comments, as everyone is trying their best to proofread the article. But authors, editors, reviewers, and proofreaders are not AI or machines; errors only make us human. Sometimes, we feel that some random person commenting on the article may not have the insight or technical expertise to evaluate the article. It seems your AI model used to write such comments is not fully operational and is making mistakes“
(Arvind R. Singh on PubPeer, June 2026)
Mohit Bajaj is a Hyper-Prolific Researcher but his choice of outlets is undiscerning. In 2024 alone he contributed 78 papers to the Scientific Reports midden – one every 5 days! – which adds up to Article Publishing Charges of over $150,000. He and his colleagues could have generated all these manuscripts, handcrafting the LLM prompts to spin their own artisanal wordwooze, but many are laden down with heavy payloads of citation payola in their References. The implication is that manuscripts were extruded from the spigots of papermills (that are also citation mills) in the manner of fancy pasta, i.e. that Bajaj outsources this task. Concerns about Bajaj’s malarkey have been sent (I am told) to the Scientific Reports Editor-in-Chief, one Rafal Marszalek – for every midden needs its own Chanticleer – inviting him to generalise beyond the three papers retracted so far (and from the ones adumbrated in PubPeer threads). But $150,000 is a strong demotivation for any sense of investigative urgency.
One Bajaj product has first priority on our attention because the title invokes EV charging and Blockchain… using a tag-team of shibboleths to lull editors and reviewers into somnolence and removing any need to blend the piffle with enough substance to make it even vaguely plausible.
Arvind R. Singh , R. Seshu Kumar, K. Reddy Madhavi , Faisal Alsaif , Mohit Bajaj, Ievgen Zaitsev “Optimizing demand response and load balancing in smart EV charging networks using AI integrated blockchain framework” Scientific Reports (2024) doi: 10.1038/s41598-024-82257-2
It appears that a Machine Learning (ML) framework was trained on a barebones 2014-2015 collection of electric vehicle (EV)-charging transactions so that it could model a Blockchain-facilitated charging market, despite the absence of any blockchain component in that archive (because 2015); or failing that, of how that might have been simulated. I hasten to add that Asenio’s original archive was not used, nor the copy at Github, but rather a pirated copy at Kaggle that subsequently evaporated. The data provenance is moot, really, as no effort was actually wasted performing the supposed ML fiddle-faddle; hence the meaningless Figure-shaped images that are shared with other customers of the same papermill.

Not to forget the pages of twaddle that are embroidered with mathematical symbols in the hope that these will make them ‘mathematical’, much as a would-be magician might embroider his robe with the glyphs of astrology and Kabalah:

Electric Vehicles are a popular topic when someone calls in papermill assistance for their ascent of the slippery pole of academia… along with battery-performance-enhancing nanotech, photovoltaic energy (with Perovskites), biomass, photothermal energy (with nanorrhea), hybrid generation, and sustainable energy. Those, one surmises, are what attract the EU research $$$.
Despite the absence of EVs I enjoyed this spicy slice of Kaggle-dependent simulation. Thank you, PLoS One!
Eman Ali Aldhahri, Abdulwahab Ali Almazroi, Monagi Hassan Alkinani, Mohammed Alqarni, Elham Abdullah Alghamdi, Nasir Ayub GNN-RMNet: Leveraging graph neural networks and GPS analytics for driver behavior and route optimization in logistics PLOS One (2025) doi: 10.1371/journal.pone.0328899
We are to believe that an anonymous logistics / delivery company based somewhere in South California decided to up their workforce-surveillance game and recorded 12000 snapshots, purportedly sampled at one-second intervals (beginning at 2023-01-01 00:00:00 and ending at 2023-01-02 09:19:59), that switch randomly between five drivers (101 to 105) driving five vehicles of an unspecified nature (1001, 2002, … 5005) without ever stopping to sleep. As well as Driver and Vehicle ID, the snapshots capture the vehicle’s location, direction, engine state, etc. Then they deposited the collection at Kaggle as ML-education material, as one abandons a surplus infant at an orphanage, so that researchers like Aldhahri et al could train their systems to recognise outlier driving and book the offenders into Retraining Camps.
Aldhahri et al never marvelled that vehicles were teleporting randomly around, so that distances of hundreds or even thousands of kilometers between separate locations recorded at consecutive seconds, while the drivers were constantly teleporting between different vehicles’ steering wheels. The “geographical distribution of trips with anomalous events” overlaps completely with normal trips, though this pales in importance when one realises that both distributions extend across North America from Nevada into the North Atlantic Ocean.

There is nothing in this ‘dataset’ that isn’t just random numbers. The authors saw nothing curious about the distribution of driving speeds, merely observing that “Readers may determine normal driving speeds, speed ranges, and outliers from this figure. The apex of the distribution indicates the most usual speeds. At the same time, the tails represent extreme driving behavior, such as extremely slow or very fast driving, which may imply unsafe or abnormal trends. This figure aids in analyzing driving behavior, which is crucial for identifying fast or erratic driving, thereby supporting the study’s emphasis on driver behavior and route anomaly detection.“
Burt! This bloke won’t Kaggle!
“This is unedifying but illustrative on so many levels of how quickly papermills have parasitised the Open Science paradigm and strip-mined all its resources. It also illustrates how quickly the incantatory Worship Words of Machine Learning shut down the critical faculties of reviewers when included in a manuscript, convincing them that its ludicrous claims are…
I previously characterised Kaggle as primarily a social-media site, with its own social-media internal attention economy of upvotes and sharing… though the approbation of one’s peers is not won by sharing memes or sick burns, but rather by uploading datasets. And just as with Manga or p0rn images, software generation is even better than tiresome observations of reality.
The source of “Driver behavior and route anomaly Dataset (DBRA24)” is Kaggle denizen “Datasetengineer“… who also donated 66 other random-number-compendia. I am confident that he or she is not addicted to that attention rush and can stop faking datasets any time. At least the choice of sobriquet gives the audience fair warning that one’s uploads were “engineered” rather than recorded from actual observations in the old-fashioned way. Anyway, another of these delirious hallucinations is a “Logistics and supply chain dataset“, which:
“…captures a comprehensive set of logistics and supply chain operations, specifically collected from a logistics network in Southern California. The data spans from January 2021 to January 2024, encompassing various aspects of transportation, warehouse management, route planning, and real-time monitoring. It includes detailed hourly records of logistics activities, reflecting conditions in urban areas and transport corridors known for high traffic and dynamic operational challenges.
The dataset is collected from various sources, such as GPS tracking systems, IoT sensors, warehouse management systems, and external data providers. It covers different transportation modes, including trucks, drones, and rail, providing insights into operational efficiency, risk factors, and service reliability. The data has been anonymized and processed to ensure privacy while preserving the information needed for analysis.”
Perhaps it is the same fictitious South Californian logistics / delivery company! Still ramping up its Workplace Surveillance infrastructure, now wiring its vehicles with sensors for an all-encompassing monitor-and-record Pantechnicon … all motivated by no obvious purpose, except of course to enrich the cognitive environment of the occupants of the Kaggle sandbox. I hesitate to ask how the intangibles of this comprehensive digital simulacrum were measured (to 16 significant digits), such as ‘driver fatigue’ – EEGs? Questionnaires every 10 minutes? Blood-sensing implants for drivers? – for fear of the whole edifice of invention crashing down to Earth when the necessary suspension of belief is disrupted.
At any rate, a certain Xiao Yang (student at Xi’an Aeronautical Institute) seized upon this farrago to add some credibility to his own pretense of training ML to … umm… do stuff.
Xiao Yang O²RDL-net for joint risk classification and delay forecasting in logistics systems using interaction Scientific Reports (2026) doi: 10.1038/s41598-026-55703-6

I am not sure what it did, despite many colourful and decorative infographics. Not that the putative model’s capability matters, any more than the aspirational nature of the analysis… for the core content of the paper is of course its payload of bogus citations to improve the academic indices of customers for the papermill’s citation-payola operation.
The Citation Payola
“The proposition that a niche of citation brokers exists, opens our eyes to other transaction options..” . Smut Clyde
I imagine that around each product of Datasetengineer‘s fertile imagination, a similar penumbra of equally irreal publications popped up, invoking it as “corroborative detail …to give artistic verisimilitude to a bald and unconvincing narrative” (tracking them down is left as an exercise for the reader). The point is that ‘made-up stuff calling upon other made-up stuff for credibility’ is emerging as a theme here. Whenever a paper informs you that a ML system was trained on a Kaggle dataset, you should think of a fly-by-night trader popping up in a Metro station to sell cheap knock-off versions of the Emperor’s New Clothes from a trestle-table stall.
By way of example, here’s Zhao Zihan promising to deliver mankind’s oldest dream – an AI career advisor. A Kaggle dataset on Education and Career Success is involved.
Zhao Zihan A multi-factor data mining and transformer-based predictive modeling approach for career success using educational and behavioral traits Scientific Reports (2025) doi: 10.1038/s41598-025-23078-9
The underlying career-advice postulate is irresistible (that for every square peg there is an equally square occupation hole waiting to be filled, so career pathways are totally not constrained by chance, prejudices, and a dysfunctional labour market); thus the failure of reviewers to click on the link comes as absolutely no surprise (also, Scientific Reports!). It would have warned them that:
“This dataset should not be regarded as serious; it was synthetically generated using real-world education and career trends. It is meant for learning purposes only.”
A functioning AI advisor would probably have counselled Z. Zihan against pursuing a career in science. I particularly liked Figure 5, presented as the output of t-SNE dimensional reduction, which is clearly an hommage to Benjamin Franklin’s “Join or Die” woodcut.

Confronted as students are with a plethora yawning absence of employment niches and under pressure to make the objectively-optimal choice, it’s no wonder that their mental health is another growth area of garbage papers. In the words of Dan Jiang: “This need urge for intelligent systems that can support mental health monitoring in educational environments.” Once again a paper is supposedly founded on Kaggle: this time a “Mental Health” spreadsheet (one of 10 donated by a Mahdi Mashayekhi):
Dan Jiang Predicting student mental health through entropy-based features and interpretable cross-attention transformer networks PLOS One (2026) doi: 10.1371/journal.pone.0347294
Jiang describes it as “2,000 student records with 10 psychological and behavioral features, such as anxiety score, depression score, productivity level, sleep hours, and physical activity. Each feature was collected via self-reported survey instruments commonly used in psychological assessment studies. The dataset includes both continuous and categorical attributes.“
In fact there are 10,000 rows in the spreadsheet – a “synthetic simulation of global mental health survey responses” – but the “employment status” variable was randomly set to “student” in 2000 of them, the rows that Jiang picked out. This is why these 2000 visitors from Random-Number-Planet have a flat age distribution in which 66- and 18-year-old tertiary students are equally well-represented. Creating characters by rolling dice is all very well for a game of Dungeons-&-Dragons but it is no way to study mental health. Even if the prevalence of Chaotic Good and Lawful Evil are your topic. Speaking of Lawful Evil, Jiang assures us that “Moreover, responsible AI practices were followed, promoting transparency and fairness in model predictions“, where “responsible AI practices” evidently include “fabulating data”.
The Men Who Stare at Ketotic Cows
“The References dwell on bovine indigestion, seizures, and the work of N. Arunkumar.” – Smut Clyde, jazz-punk Klezmer musician.
For the sake of variety, meet Huang & Jiang, who made up their own artisanal hipster fake data rather than outsource that task to the unreliable provenance of Kaggle.
Xianwei Huang , Wei Jiang Utilizing multi-level convolutional neural networks to achieve refined modeling and visual analysis of college students’ mental health data PLOS One (2025) doi: 10.1371/journal.pone.0335048
Allegedly, 500 students Sims characters devoted so much time to evaluating themselves on “widely used, standardized scales for assessing mental health outcomes such as anxiety, depression, and stress. Specifically, the Generalized Anxiety Disorder Scale (GAD-7), Patient Health Questionnaire-9 (PHQ-9) for depression, and Perceived Stress Scale (PSS)” that they had no time left for studying. Meanwhile “participants’ responses were reviewed by mental health professionals, including trained counselors and psychologists, who provided additional judgment and context regarding the severity of the reported symptoms. This professional review helped ensure that the self-reported data were interpreted correctly“. This fine-grained informational firehose boiled down to Table 1:

All this comes from the “mental health assessment system” of a totally not fake university, “which also provides a virtual community data visualization function (see Fig 2)“. Figure 2 (right) is either a “virtual community data visualization function” in the Cyberpunk stylings of Neuromancer, or an imagined university campus depicted by real-estate software.
But N=500 is too small so (as is the custom in the ML literature) the authors supplemented their real made-up cases with additional fake made-up cases:
“… we used data augmentation techniques such as rotation, translation, and scaling to expand the training set and improve the model’s generalization ability.”
The reviewers did not wonder how you “rotate, translate and scale” mental health scales. Fig 8 purports to show how “data augmentation provides a modest but consistent improvement in accuracy”. Personally I believe it to be a Seurat painting of a Cadbury’s Easter Egg.
Research on Intelligent Trash Can Garbage Classification Scheme
“It’s not as if the Special-Issue Guest Editors or the imaginary ‘Peer Reviewers’ pay attention to the provenance of the images that fill the Figure-shaped gaps, or care whether the supposed alternatives in these horse-races are even algorithms at all.” – Smut Clyde
“What those papers are missing,” you may think, “is the element of intrusive round-the-clock surveillance”. Certainly Musthafa et al thought that way, along with “more IoT and enhanced security please!” They may have been inspired by their Ref. [19], which describes how students in a class on IoT Security were shown how to create personal fake-databases to practice on:
A. Syed Musthafa, G. Jenifa, Lavanya Murugan, Usha Moorthy IoT Based Security Enhanced Continuous Health Monitoring Framework Using Lightweight Cryptography and Deep Learning Techniques International Journal of Computational Intelligence Systems (2026) doi: 10.1007/s44196-026-01237–
“We used immediate student dataset [19] comprising 1100 instances to evaluate student healthcare.”
[19] Rao, A.R., Elias-Medina, A.: Designing an internet of things laboratory to improve student understanding of secure IoT systems. Internet Things Cyber-Physical Syst. 4, 154–166 (2024)
Alas, apart from fleeting references to “certain blood pressure hours” and “daily sleep hours”, Musthafa et al say nothing about this “publicly available real time student dataset” they fabricated… they content themselves with claiming that their poorly-characterised ML system recaptured its underlying structure better than equally-uncharacterised alternatives. Nothing about this hand-waving performance of impoverished imagination stands out to attract our attention… but it gets better with Fig 8!

“Kaggle Elikplim dataset”? Though uncredited, this is an Energy Efficiency Database – one of eight deposited by Kaggle donor “Elikplim” (Ahiale Darlington). It consists of energy efficiency values simulated in Ecotect software for 768 situations (12 basic buildings, with 64 variations of the parameters). Analysed with health-monitoring software. This is sorcery.
Shake the Stupid Tree and see what falls out
“Does this mean it’s time for an update on the bogus-citation economy? Leonid thought it is, and now you all must suffer for his misdirected priorities. ” – Smut Clyde
Moving right along, it is time that the readers learned about the Bonn EEG dataset. Andrzejak et al 2001 report how they fitted five patients with intractable Temporal-Lobe Epilepsy with cerebrally-implanted electrodes for round-the-clock recording of their temporal-lobe activity. 300 segments were snipped out of the footage (each 23.6 seconds long) and filed in three folders of 100 pre-seizure, 100 during-seizure, and 100 post-seizure segments. Andrzejak et al. also recruited five non-epileptic volunteers and recorded 200 more samples to fill two more folders (because reasons)… using external electrodes, meaning that comparisons are not meaningful. A pox upon those tiresome ‘Review Boards’ and their ‘ethics approvals’!

Clearly it would be stupid and pointless to train a ML system on those internal-electrode recordings and expect it to recognise the onset of a seizure in external signals. Which is to say, there is a huge literature of make-believe research where authors pretend to have done just that.
Wail Mardini, Muneer Masadeh Bani Yassein, Rana Al-Rawashdeh, Shadi Aljawarneh, Yaser Khamayseh, Omar Meqdadi Enhanced Detection of Epileptic Seizure Using EEG Signals in Combination With Machine Learning Classifiers IEEE Access (2020) doi: 10.1109/access.2020.2970012
Since 2001 the Bonn recordings have been rehoused many times as links decay, but few bother to find a current link to access them (or to RTF Paper as the link scoldingly advises). Instead they claim to use a bootleg copy from 2017 in the UCI Machine-Learning Archive. This no longer exists… taken down, I suppose, due to the stolen-IP issue.

But wait, it gets stupider… in fact the UCI version consists of 11500 brief clips that Wu and Fokoue from Rochester Institute of Technology created by dicing each segment into 23 1-second sub-segments. Wu mirrored these on GitHub, then a certain Harun-Ur-Rashid copy-pasted them to Kaggle without crediting Wu or Fokoue; and this now-vanished meta-piracy hoard is the dataset that EEG bullshit artists pretend to use in the disposable papers that swarm and pullulate in IEEE journals and beyond. Notably, after two generations of cumulative stupidity, Ur-Rashid’s explanation is so engarbled that
- Everyone believes that they came from a representative population of 500 subjects, not just 10. Contamination issues from training and testing on the same subjects don’t apply.
- Many fraudsters are convinced that they’re not dealing with waveforms at all (each with 178 time-points), but rather with lists of 178 acoustic features abstracted from the original waveforms in some unspecified way (and therefore comparable across the clips).
Despite this profound and pervasive incomprehension, everyone’s analyses always work – diagnosing seizures with 100% accuracy that is better than everyone else’s 100% accuracy. A few examples will suffice.
- Sanaa Al-Marzouki Advancing epileptic seizure recognition through bidirectional LSTM networks Frontiers in Computational Neuroscience (2025) doi: 10.3389/fncom.2025.1668358
- Ayut Ghosh, Arka Prava Roy, Ramapati Patra, Hemanta Kumar Mondal Designing Efficient NoC-Based Neural Network Architectures for Identification of Epileptic Seizure SN Computer Science (2021) doi: 10.1007/s42979-021-00756-9
- Muhammad Jahanzeb, Abdul Hannan Khan, Shakeel Ahmed, Abdulaziz Alhumam, Muhammad Farrukh Khan, Shahan Yamin Siddiqui Privacy preserving epileptic seizure recognition using federated and explainable machine learning Discover Computing (2026) doi: 10.1007/s10791-026-09956-4
A second corpus of epilepsy-EEG recordings exists, collected at the American University of Beirut and uploaded by Wassim Nasreddine to Mendeley. A certain Omid Mahdiyar (PubPeer record) purported to base his papers on these rather than on the Bonn EEGs, but it would be a shame to omit them here, as there is a kind of grandeur in the scale of their incompetent mendacity.
- Mahdiyeh Lak, Jasem Jamali, Nahid Adlband, Mehdi Taghizadeh, Omid Mahdiyar Hybrid fuzzy machine learning models optimized with meta-heuristics for accurate EEG-based neurological assessment Scientific Reports (2026) doi: 10.1038/s41598-026-35669-1
- Solmaz Badr, Jasem Jamali, Mehdi Taghizadeh, Nahid Adlband, Omid Mahdiyar EEG based epileptic seizure detection using SVM fuzzy learning and metaheuristic optimization Scientific Reports (2025) doi: 10.1038/s41598-025-24431-8
The former paper has Figure 1 (“EEG signal changes in different brain states“), which I choose to interpret as spectrograms of me reciting the Hedgehog Song as the intake of beer continues; they are certainly not EEGs. The Grok logo at lower right suggests ‘AI fabrication’. Not to be outdone, the latter paper gives us Figure 10 – another of those t-SNE Easter-eggs in which two distributions belie the authors’ claim that they are distinct, by overlapping completely.

Both purloin their descriptions of the EEG corpus, and diagrams, from Vieira et al 2023. The engarbaging paraphraser software somehow turned ‘ictal’ into the novel acronym ‘CATAL’.

The former paper is arguably the absurder of the two. The authors declined to be commit themselves to the purpose of the exercise in ML: whether it was to recognise a seizure, or to determine a comatose patient’s level of consciousness.
They also vacillated about which menagerie-themed optimiser algorithms they used, announcing these as “Goose and Grey Wolf” strategies while going on to explain a Starfish and a Watercycle algorithm. So really there are four shite papers here, condensed into one; no wonder the Scientific Reports reviewers accepted it!

Not to forget the novelty of Shannon Violet’s Entropy. And Table 1, supposedly explaining the three panels of Figure 1 in terms of 5 classes of EEG, all harking back to Andrzejak et al 2001 – so there is an excuse to include Mahdiyar’s fiddle-faddle after all.

Badr et al 2025 do in fact explain The Goose Optimiser. For certain values of “explain”. There may or may not be bizarre metaphors. Geese may or may not store rocks in their feet.

Bottom of the barrel: BatDolphin-based sparse fuzzy algorithm
“BatDolphin-based sparse fuzzy algorithm, cat swarm optimization, honey bees optimization, moth amalgamated elephant herding optimization, fitness sorted moth search algorithm, improved tunicate swarm optimization, lion algorithm, deer hunting optimization, various rider optimization schemes, grey wolf optimization, cuckoo search, and finally a bat algorithm. Such a zoo of names immediately raises suspicion, and for a good…
Still on the neurology theme, that fleeting reference to spectrograms brings Parkinson’s Disease (PD) to mind. For PD affects speech, introducing jitter and shimmer and other spectrographic subtleties as the loss of dopaminergic cells degrades fine muscle control in the larynx, and inevitably there is a public Oxford PD Onset Detection Dataset. A table with 197 rows – speech samples from 23 PD patients and 8 healthy controls – and 23 acoustic qualities. For some reason this is housed at the UCI Repository; the entry also includes a bonus file from a different project on Telemonitoring by the same authors, with only 16 spectroscopy features for 42 unrelated people (all with PD), collectively pronouncing the sustained vowel /a/ for 5875 times in a “Sound like a Sheep” competition. Keep all of this in mind. There will be a test.
The Kaggle Dance was performed. Soon Alhawiti 2025 rocked up in Sensors, claiming to have accessed two PD datasets: “Two publicly available datasets were utilized to construct the multi-modal fusion pipeline”. The description of the first identifies it as the Oxford set from [41] – somehow attributing it to Sakar et al. [42])
Khaled M. Alhawiti Multi-Modal Decentralized Hybrid Learning for Early Parkinson’s Detection Using Voice Biomarkers and Contrastive Speech Embeddings Sensors (2025) doi: 10.3390/s25226959
“UCI Parkinson’s Dataset: The proposed study used the Parkinson’s Disease Classification Dataset available from the UCI Machine Learning Repository (ID 470), originally developed by Sakar et al. [41,42]. The dataset contains 195 sustained-phonation voice recordings collected from 31 participants (23 diagnosed with Parkinson’s disease and 8 healthy controls). Each sample includes 22 handcrafted acoustic biomarkers such as jitter (local, RAP), shimmer, fundamental frequency, harmonics-to-noise ratio, and mel-frequency cepstral coefficients”

Our interest, though, lies in Alhwiti’s second source of a Machine-Learning curriculum:
“DAIC-WOZ Corpus [43,44]: This clinical dataset includes high-quality WAV recordings (mono, 16 kHz) of structured psychological interviews. Voice samples were denoised via spectral gating and segmented using energy-based voice activity detection (VAD). Each utterance was then passed through a self-supervised speech encoder (HuBERT or Wav2Vec 2.0) trained with a SimCLR contrastive framework, producing a 768-dimensional latent embedding per recording.”
The DAIC-WOZ corpus is not publicly available. The on-line repository “DAIC-WOZ Database” provides a webpage where one can apply for access, but Alwahiti didn’t claim to have done so. Nor did he explain how speech recordings and transcripts from one group of distressed & depressed interviewees could be useful for training an ML system to detect PD in spectroscopy features from an unrelated group of speakers. Sorcery!!
The Oxford dataset also features the following study from India, those authors didn’t bother to provide a link, or to attribute Little et al. (2007) as the user license requires, but Scientific Reports so no-one cared. As with everything examined here, the paper is primarily a citation-delivery vehicle. For instance, a stack of citations for the benefit of some E. Aslan, added at a late proof stage.
Pradeepta Kumar Sarangi , Rajnish Srivastava, Monica Dutta, Ashwin Dobariya, Subhanshu Goyal , Sunil Lavadiya, Samah Alshathri , Walid El-Shafai Enhancing prediction accuracy for Parkinson’s disease using advanced machine learning models Scientific Reports (2026) doi: 10.1038/s41598-026-54057-3

We are really only here to admire Figure 1:

Tag yourself! I’m Matemal nurfurnrg!
Crunchy Frog and Cockroach Cluster
“On one side: late-career scientists resorting to purchased promotion of their early-career papers. On the other side: whole new genres of paper-shaped artifacts that are little more than packaging for ever-larger citation cargoes, and papermillers no longer bothering to find buyers for naming rights on their products.” – Smut Clyde
With all that out of the way, at last I can circle back to Mohit Bajaj. For he is an old friend of PubPeer; we find him in frequent collaborations with a Ukrainian colleague – Ievgen Zaitsev of Kyiv (PubPeer record). And especially Vojtěch Blažek and others (Lukas Prokop, Milkias Berhanu Tuka (PubPeer records here, here and here): a circle centred on Ostrava Technical University in Czechia – which is not unknown to For Better Science readers. Indeed, so deep is Bajaj’s patriotic identification with the Czech Lands, as Corresponding Author he often uses his ‘mb.czechia.gmail.com‘ e-account. He is popular and much-run-after as coauthor because his specialty discipline in electrical engineering lends itself well to the publishing imperatives discussed above.
I, Rajender Varma, Highly Cited Researcher
“I could not comprehend the situation where a university picks up on individuals with an extraordinary and sterling performance and basically destroy one of the top European institutions. ” – Raj Varma
Readers who pay attention will recall that we began with a Bajaj / Zaitsev collaboration, illustrated with nonsensical and garish Figure-shaped images – which it shared with another paper. That turns out to be this one:
P. Santhiya, Kogilavani Shanmugavadivel, R. Rajalakshmi, N. Krishnamoorthy Cross-modal Synergy for Enhancing Emotion Recognition Through Integrated Audio–Video Fusion Techniques International Journal of Computational Intelligence Systems (2025) doi: 10.1007/s44196-025-00811-w
Sylvain Bernès noted in the PubPeer thread that the mathematical argumentation in that paper is gibberish – just another torrent of random sigils.

But here we are concerned with the tequila hangovers Figures. Just look at them!

In the long-standing Marriage-in-Cana tradition I have saved the best wine for last! “The best” in this case being another document signed by Anas Bilal (obvious anagram), who featured in the previous post. I do not want him to feel left out here. Here he informs us that
“It is necessary to identify the problem, collect the data, and pre-process it with the help of credible sources, such as Kaggle.”
Anas Bilal, Waeal J. Obidallah, Sobia Wassan, Mubarak Albathan, Riyad Almakki, Zeyad Alshaikh, Muhammad Shafiq Fusion of genomic and pathological data for breast cancer detection using BCDNN Frontiers in Medicine (2026) doi: 10.3389/fmed.2026.1726223
Figure 3 is a special delight:

Yes, these are digital representations, in the sense that fingers were used to paint them. But it would be churlish to complain, considering that FNA is never mentioned again.
The title of the paper is unambiguous: Bilal and his coauthors fine-tuned a neural network so that genomic data and histopathologic images will be combined when they flow into the hopper of the sausage-machine, thereby classifying a tumour as benign or malignant. They reiterate this claim throughout the text…
- “The number of neurons in the input layer should be equal to the number of features obtained from the genomics and histopathological data.”
- “Due to this gap, the current paper examines deep learning-based integration of genomic and histopathological data to diagnose breast cancer better.”
- “This paper presents the suggestion of a BCDNN, which combines both genomic and histopathological data in a single learning framework.”
- “In comparison, BCDNN combines both the features of genomics and histopathology conditions, which allows the model to learn the complementary data on the molecular and tissue levels.”
- “A major novelty of this work is the combination of both genomic and histopathological data, which is directly related to the enhanced performance.”
- “This study demonstrates the potential of the proposed BCDNN for supporting breast cancer classification using genomic and histopathological features.”
… interrupted only by the Fine Needle Aspiration, and by Table 8 (which seems to have wandered in from some other paper, as it boasts of the diagnostic superiority of the authors’ neural net for Dynamic Thermographic scans).

Then we encounter “A publicly available breast cancer dataset from Kaggle… encompassing both genomic and pathological features.” This turns out to be the Wisconsin Breast Cancer Dataset, which indeed has been copy-pasted to Kaggle… sadly, it includes absolutely no genetic or genomic data. I can only suppose that the performance values presented by the authors are aspirated aspirational Artist’s Impressions of the values they hope for, were their proposed system to be (a) implemented and (b) applied to a hypothetical collection of more appropriate data.
This is not the stupidest paper ever published in Frontiers, not by far, but nor is it the publisher’s finest display of rigorous editing and peer-reviewing.
Another Bilal paper is all about “Quantum computational infusion”, which is when Schrödinger may or may not have included a teabag in your cup of hot water. Thus in Figure 2a a mammograph is shown twice, simultaneously benign and malignant: we won’t know which until the wavefunction collapses.
Anas Bilal, Muhammad Shafiq, Waeal J. Obidallah, Yousef A. Alduraywish, Haixia Long Quantum computational infusion in extreme learning machines for early multi-cancer detection Journal of Big Data (2025) doi: 10.1186/s40537-024-01050-0

Though Figure 2d is sourced from the Lung Image Database Consortium according to the legend, it shows two successive coronal-section scans of the same head… again, the quantum wavefunction has not collapsed so the parietal-lobe meningioma is benign and malignant.

However, what really commends the paper is its immodest self-congratulatory tone. In parallel with assembling the text from ChatGPT responses, the authors were prompting ChatGPT to write gloriously positive reviews for later, but included them in the manuscript by mistake.
“Developing the Q-GBGWO-ELM model is a crucial event in medical diagnostics, as it contributes to progress in the early detection of multi-cancer […] The Q-GBGWO-ELM model stands out beyond other methods of feature extractions and parameterizing ELM, as it has allowed for achieving record-breaking levels of diagnosis accuracy.[…] It is a versatile and dependable tool that will cement predictive healthcare in the future. “

Evidently the authors’ diagnostic breakthrough is already successfully in use, though only by “Mexican clinicians and researchers”. It’s unmitigated nonsense but the reviews were positive.

Donate to Smut Clyde!
If you liked Smut Clyde’s work, you can leave here a small tip of 10 NZD (USD 7). Or several of small tips, just increase the amount as you like (2x=NZD 20; 5x=NZD 50). Your donation will go straight to Smut Clyde’s beer fund.
NZ$10.00


0 comments on “The subtleties a spectrograph would miss”