The key intended audience of these sites is not concerned Americans, it’s not even humans—most of the sites average a few hundred unique visitors each month. Instead, Parscale and his firm, Clock Tower X, created them as part of a $46.5 million contract with the Israeli government to try and influence artificial intelligence-powered chatbots, tools like Claude or ChatGPT.

While this is a drop in the bucket of Common Crawl data, *emerging research suggests a very low number of documents can successfully manipulate AI chatbots, no matter the data volume.

I shared this article specifically for the above mechanism described as it is a newish vector of disinformation/misinformation.

*https://pmc.ncbi.nlm.nih.gov/articles/PMC12881903/

Health care artificial intelligence (AI) systems now play a significant role in influencing diagnosis, documentation, triage, treatment planning, and resource allocation. As adoption accelerates, these systems face growing exposure to data poisoning attacks that can subtly and systematically degrade model performance. Even small adversarial manipulations can propagate across clinical workflows and affect large patient populations before they are detected. Consider a representative scenario: a radiology AI deployed across a hospital network begins missing early-stage lung cancers disproportionately among specific demographic groups. The errors resemble known health care disparities and therefore do not raise an immediate alarm. Yet, the root cause is a small set of approximately 250 poisoned images—comprising only 0.025% of a million-image training dataset—inserted during routine data contributions by an insider. Detection occurs years later through retrospective epidemiological review, long after patients have experienced delayed diagnoses and poorer outcomes.

This hypothetical case reflects empirically demonstrated vulnerabilities. Recent security studies have shown that health care AI systems can be backdoored with as few as 100-500 poisoned samples, ->->->regardless of total dataset size<-<-<- [1-5]. Attack feasibility has been confirmed across several architectures, including large language models (LLMs) used for clinical documentation and decision support [1], convolutional neural networks (CNNs) used in radiology and pathology [3], and emerging agentic systems that autonomously assist with clinical tasks [6]. These attacks do not require privileged system access; routine insider access to data-collection workflows is often sufficient [1-4]. A counterintuitive but critical finding from recent security research is that successful poisoning attacks require only 100-500 malicious samples, independent of total dataset size [5]. This challenges the conventional assumption that scaling training data provides security through dilution and has profound implications for health care AI, where training datasets routinely contain millions of samples yet remain vulnerable to attacks from a single insider over weeks or months.

  • GardenGeek@europe.pub
    link
    fedilink
    arrow-up
    9
    ·
    4 days ago

    The wow-effect during the initial introduction of LLMs was caused by their ability to give useful answers to not-to-specific questions by guessing words. If thei data is manipulated by a variety of players they not only are insanely resource intensive (at least the large models) and, by design, prone to hallucinate due to stochadtics but also become useless as their ability to answer correctly is imperiled by a blurred data basis. If models start answering simple questions like 2 plus 2 with 5 because they’re being manipulated by big5 they’re simply not useful anymore.