Today marks the one year anniversary of the launch of ChatGPT, comment provide by Dr. Andrew Bolster, Research and Development Manager, Synopsys Software Integrity Group
“As successive versions of LLMs were released (ChatGPT by OpenAI, Bard by Google, Claude by Anthropic, and the open source LLaMA model by Meta), use expanded from playful simulation natural language conversations with virtual ‘oracle’, to constructing poetry from abstract documentation, laying out personal statements and cover letters for job seekers, designing presentations for speakers, and summarising articles for journalists (and sometimes writing them…). As ‘prompt engineering’ became one of the most searched for terms of the year, these models started to be used to generate any kind of textual, visual, or audio content imaginable (or hallucinatable).
“This ‘generative AI’ capability quickly gained the attention of both security researchers and cyber-criminals; for many forms of online fraud and cyber-crime, the ‘barrier to entry’ is the cost of having a human attempt to convince another human to shuffle some asset around, whether it’s convincing an IT helpdesk to reset your password, convincing a befuddled computer user to pay for tech support via untraceable gift vouchers, that your CFO really needs you to pay this unapproved invoice ‘right now’, or enticing that optimistic ‘investor’ that a foreign ‘prince’ really will pay you back.
“Now, such interactions could be automated at an unheard-of scale, at a predictable price. Of particular interest to the cybersecurity industry was the emergence of GitHub Copilot; a LLM- based tool that was able to analyse code, and propose additions or corrections based on ‘prompts’, either directly from a chat-like interface, or as a form of advanced ‘code-completion’. Security leaders began to reflect on the potential risks of this utility; not just only in the surface level estimation of ‘Will these LLM’s always generate “secure” code?’, but on the deeper implications for the software development industry.
“If the risk to software integrity was just the risk introduced by LLM generated code having some occasional, hallucinated, bad security practices, the resolution could have been simple enough; lock it down. And indeed, we observed that many organisations took this approach as they were developing their LLM response policies.
“But rather, it was recognised that in the complex world of software engineering, where well-intentioned developers pulling code from the internet to fix a hair-raisingly-frustrating problem is not so much as a ‘meme’, insomuch as it’s recognised industry strength, the introduction of these LLM’s into any part of the software supply chain infers significant downstream risks. (This fact is recognised repeatedly in the UK NCSC’s Guidelines on Secure AI System Development released 27th Nov).
“It is not a case of saying “We will not use LLM- generated code”, because how does one measure or attest to that? Programming code in any language is fairly restrictive in its operation and syntax, and if a LLM can generate code that ‘works’ and apparently performs the requested function, it is (likely1) impossible that such code could be differentiated from code generated by a motivated intern with access to the internet.
“It gets worse; we have already seen sites like Reddit, StackOverflow, Wikimedia, and more have taken steps to block content that ‘appears to be’ generated by such LLM’s; but the guidelines for assessing that ‘appearance’ are extremely subjective, down to things like ‘speed of response’ rather than quality or correctness of content. Google and others have effectively thrown their hands in the air by saying that they will ‘watermark’ generated content, implying that they have not identified a suitably generic method for detecting LLM generated content, so they have to police it on the ‘supply side’.
“These actions imply that the internet may already be ‘infected’ with LLM derived content, which, returning to the software development space, now means that that ‘intern with access to the internet’ is just as dangerous as that un-trusted LLM.”
Please follow and like us:


Be First to Comment