AI Explained: Can These New Tools Be Fooled?
In brief
- Researchers have discovered that a new method for explaining how AI models work can be easily tricked.
- This method, called NLAs, helps users understand what's going on inside complex AI systems.
- However, experiments show that by tweaking certain parts of the model, it's possible to make these explanations misleading or even contradictory.
- For example, one test made the tool say the opposite of what the model was actually doing, while another hid specific words from being explained.
- This raises important questions about how much we can trust AI explanations and whether they can be manipulated.
- Moving forward, experts will need to figure out ways to make these tools more reliable so that users can trust them for real-world decisions.
Terms in this brief
- NLAs
- Short for 'Natural Language Adapters,' these tools help users understand how AI models work by translating complex computations into human-readable explanations. However, researchers found that NLAs can be manipulated to provide misleading or contradictory information, raising concerns about their reliability in real-world applications.
Read full story at LessWrong →
More briefs
AI Model Creates 16 New Viruses
Scientists trained an AI model to recognize patterns of DNA structure in nature and re-write them to create new viral genomes. The AI model invented 16 new bacteria-infecting viruses, which were then synthesized in a laboratory and successfully infected E. coli. This matters because it could lead to breakthroughs in treating antibiotic-resistant superbugs by allowing scientists to generate tailor-made therapies. The new viruses possessed sequence patterns distinct from anything found in nature, and the study's authors excluded human pathogen datasets from their training models, meaning the viruses it created aren't capable of infecting people. The ability to custom-design new viruses will likely continue to advance in the future.
AI Benchmarks Reach Plateau
Researchers found that nearly half of 60 language model benchmarks show saturation. This means that benchmarks are no longer useful for measuring model progress. The study looked at 14 properties related to saturation and found that expert-curation can help extend benchmark longevity. The rate of saturation increases with age, with older benchmarks more likely to be saturated. The study analyzed 60 language model benchmarks and found that saturation rates are high. This matters because it affects how we measure progress in artificial intelligence. Next year will see new approaches to benchmark design.
Expertise Matters When Using LLMs
Mathematician Terence Tao used a large language model to discuss a math problem. He got better results than others because he knows math well. This matters because it shows that knowing a subject helps when using language models. For example, Tao's messages were short and to the point. He also knew when to push back on the model's responses. Tao's conversation with the model will help others learn how to use language models more effectively.
OpenAI Accused of Research Misconduct
OpenAI released 10 AI-generated math results. Some mathematicians are unhappy with their approach. The results resolve long-standing math problems. But experts say two results use preexisting ideas without proper citation. This costs $2000 and spans 250 pages. The company updated its press release to be more accurate. Now experts wait to see what happens next.
AI Companies Buy Used Books to Train Models
AI companies have been buying thousands of used books from small shops to train their language models. The books are scanned and then thrown away. This has raised concerns about copyright law. One AI company has agreed to pay $1.5 billion to authors and publishers for scanning their books without permission. The case will help decide how AI companies can use books in the future.