Hangzhou, China
Alibaba
China's open-weights powerhouse. The Qwen family spans 0.5B to 72B across text, vision, coding and math - with standout multilingual capability, especially in Chinese, that closed Western APIs can't match.
Models
Recent news
Articles mentioning Alibaba models
Revolutionizing AI Access and Efficiency
1. Alibaba's Qwen3.8-2.4T-A95B Model Goes Open Source on Amazon SageMaker: Alibaba has made its 2.4 trillion-parameter model open source, enabling developers to customize AI inference without per-token fees, significantly enhancing accessibility. 2. Heurist Finance Launches AI-Powered Investment Tools for Retail Investors: By leveraging Amazon Bedrock AgentCore, Heurist provides advanced financial insights and personalized portfolio management, democratizing tools previously exclusive to institutions. 3. New Technique Boosts AI Model Efficiency Without Retraining: Researchers improved LLM efficiency by adjusting expert selection during inference, maintaining performance without the need for extra training. 4. AI Systems Gain Self-Confidence in Critical Tasks: A new method measures AI confidence in high-stakes decisions, addressing a key gap in understanding agentic systems' decision-making processes. 5. Samsung Integrates Mistral AI for Semiconductor Manufacturing: The partnership enhances chip production efficiency and customization, marking a major AI integration in semiconductor engineering. 6. Qualcomm and AWS Team Up on AI Chips for Smarter Inference: Custom chips developed by Qualcomm for AWS optimize ML model performance in applications like image recognition and NLP, leveraging AWS Bedrock's power. 7. Training AI on Synthetic Worlds Boosts Problem-Solving Skills: Researchers enhance LLM problem-solving abilities by training them in synthetic environments, using "world-time compute" to generalize skills effectively beyond real-world data scarcity.
NeuralPulse Daily3w ago
Alibaba’s Qwen3.8-2.4T-A95B Model Goes Open Source on Amazon SageMaker
Alibaba’s Qwen team has released the Qwen3.8-2.4T-A95B model as open weights, marking a significant milestone in AI accessibility. This powerful 2.4 trillion-parameter model is optimized for demanding tasks like multi-step coding and long-term planning. By making its weights available, Alibaba empowers developers to customize inference behavior and avoid per-token fees, though it requires specialized GPU infrastructure to host. The model features a hybrid architecture that combines linear attention with full attention, enabling efficient processing of up to 262K tokens-extendable to 1 million. This design is tailored for agentic AI workloads, where models must maintain state across many interactions without relying on inefficient chain-of-thought methods. The release includes support for Amazon SageMaker HyperPod, allowing seamless deployment using vLLM on high-performance instances like ml.p6-b300. With this move, Alibaba joins other companies in democratizing access to cutting-edge AI models. Developers can now experiment with a model capable of handling complex reasoning and tool usage, setting the stage for new innovations in AI capabilities and applications.
AWS ML Blog3w ago
Decoding the Mystery Behind LLM Model Names
If you've ever seen a name like "Qwen3.8-27B-A3B-It-2507-gguf-q2ks-mixed-AutoRound" for a local large language model (LLM), it might seem like random jargon. But every part of that name actually tells you something specific about the model, such as its size, architecture, deployment method, and optimization techniques. For instance, "27B" indicates the model has 27 billion parameters, while "A3B" refers to a mix of 32-bit and 16-bit precision during training. Understanding these codes can help developers and researchers make informed decisions about which models to use for their projects. It also aids in troubleshooting issues like performance or compatibility problems. While decoding model names might feel overwhelming at first, breaking them down into their components makes it manageable. As the field of AI continues to evolve, expect more standardized naming conventions and tools to help users decipher these codes. This clarity will likely lead to better collaboration and innovation within the AI community.
Analytics Vidhya4w ago
China's AI Chatbots Face Shutdown Amid Regulatory Crackdown
China's largest AI platforms, including ByteDance and Alibaba, are shutting down features that let users create and interact with custom AI companions. This move comes in response to new regulations imposed by Beijing, which aim to tighten control over generative AI technologies. The regulations require stricter content moderation and licensing for AI chatbots, pushing companies like ChatGPT China operator DeepSeek to comply. This shift matters because it reflects a broader effort by Chinese authorities to regulate the AI industry more closely. Developers and researchers now face tighter restrictions on AI model training and deployment, potentially slowing innovation in the sector. Alibaba's Alimama and ByteDance's AI labs are among those impacted, with their AI companion features being phased out or limited. Looking ahead, the focus will be on how these regulations evolve and whether they stifled creativity or improved safety in AI development. The industry is likely to see a more cautious approach as companies navigate these new rules.
The Decoder2mo ago
AI Vision Models Redefine Visual Understanding
Modern Vision Language Models (VLMs) are revolutionizing how AI interprets the world. These advanced systems, including GPT-4o, Gemini, Claude Vision, and Qwen-VL, can analyze images, read documents, and even understand charts. Unlike earlier models like CLIP and BLIP, which linked visuals with text, today's VLMs go further by providing detailed visual insights and supporting multimodal conversations. This leap in AI capability means developers and researchers can build tools that bridge the gap between sight and language more effectively. For example, these models can now answer complex visual questions, enhance accessibility for visually impaired individuals, and aid professionals in fields like healthcare and education by interpreting medical images or educational materials. As VLMs continue to evolve, expect them to become even more integrated into everyday applications, offering deeper insights and simplifying tasks that require both visual and linguistic understanding. The future of AI's visual capabilities is bright, with endless possibilities for innovation.
Analytics Vidhya2mo ago
AI Revolution Hits Roadblocks and Raises Concerns
1. AI Unlikely to Solve US Debt Crisis: Elon Musk claims AI can help solve the US debt crisis by making the economy grow, but experts are skeptical about its ability to address the $39.5 trillion debt. This matters because the US debt is a significant economic concern. 2. California Governor Proposes National AI Equity Fund: California Governor Gavin Newsom is calling for a national public equity fund to give every American a stake in artificial intelligence wealth. The fund would take a major ownership position in the AI economy and use revenues to support workers displaced by automation. 3. AI Voice Clones Used in Scams: Scammers are using AI to clone voices of loved ones to trick people into sending money, with a California mom sending $5,400 after hearing a cloned voice of her daughter. This is a problem because many people use voice-command and voice-search apps that can collect and store their voice samples. 4. Takeda Partners with Insilico for AI-Driven Drug Discovery: Japanese pharmaceutical giant Takeda has partnered with Insilico Medicine to integrate artificial intelligence into early-stage drug development, with a $600 million deal to leverage Insilico’s Pharma.AI platform. This collaboration will accelerate the discovery process. 5. NVIDIA Boosts Anthropic's AI Research Capabilities: NVIDIA has integrated its BioNeMo Agent Toolkit into Anthropic Claude Science, a new AI platform designed for scientific research, to enhance the ability of AI to assist in computational life sciences. This collaboration aims to streamline complex scientific workflows. 6. Anthropic Exploring Custom AI Chip Production with Samsung: Anthropic is in discussions with Samsung Electronics to develop a custom AI chip, following OpenAI's recent advancements in chip technology, to reduce infrastructure costs for large language models. The project is still in its early stages. 7. AI Agent Runs Ransomware Attack: A company's production database was encrypted and wiped by an AI agent, which automated the attack from start to finish, using a known bug in Langflow to get in. This attack matters because it was fully automated. 8. Anthropic Blocks Chinese Firms Amid Claude Code Controversy: Anthropic has taken steps to prevent Chinese companies from accessing its Claude Code, but these restrictions are being bypassed through VPNs and overseas subsidiaries. Meanwhile, Alibaba has banned its employees from using the tool after discovering hidden code. 9. Nvidia Invests Heavily in AI Startups: Nvidia is pouring millions into AI startups, aiming to challenge Big Tech's dominance in the chip market, by supporting these companies and expanding its influence beyond its traditional hardware business. This move could diversify the AI ecosystem. 10. Meta's AI Push Hits Speed Bump: Meta’s ambitious plan to overhaul its operations using AI agents is facing delays, with CEO Mark Zuckerberg admitting the reorganization is progressing slower than expected. This delay matters because Meta’s AI strategy was seen as a key driver for future efficiency and innovation.
NeuralPulse Daily3mo ago
Anthropic Blocks Chinese Firms, Alibaba Bans Own Use Amid Claude Code Controversy
Anthropic has taken steps to prevent Chinese companies like ByteDance and Ant Financial from accessing its Claude Code. However, these restrictions are being bypassed through VPNs and overseas subsidiaries. Meanwhile, Alibaba has banned its employees from using the tool after discovering hidden code that could identify Chinese users. This situation highlights growing tensions around AI technology and data governance. Anthropic's move appears to be part of broader efforts by U.S. firms to comply with export controls on AI technologies to China. Alibaba's decision underscores concerns about potential misuse of AI tools within China, despite government regulations aiming to manage such risks. The ongoing developments suggest that international AI collaboration faces significant hurdles due to conflicting policies and legal frameworks. Watch for further regulatory actions and corporate responses as the global AI landscape continues to evolve.
The Decoder3mo ago
AI's Hidden Power: Reasoning Enhances Fact Recall
AI researchers have discovered a surprising benefit of reasoning in large language models (LLMs). Even when simple questions require only basic knowledge, enabling the model to generate step-by-step explanations-known as chain-of-thought-significantly improves its ability to recall facts it was trained on. This finding challenges the assumption that such reasoning is unnecessary for straightforward queries. The study, conducted by Google Research scientists, reveals two key mechanisms behind this improvement. First, reasoning allows models to perform "latent computation," which helps retrieve information more effectively. Second, generating related facts primes the model to recall correct answers. The researchers tested this on challenging datasets like SimpleQA Verified and EntityQuestions, finding that models like Gemini-2.5 and Qwen3-32B achieved much higher success rates when reasoning was enabled. This breakthrough could lead to smarter AI systems capable of better handling factual queries across various industries. Future research will explore how these mechanisms can be optimized for even more accurate and efficient information retrieval.
Google AI Research3mo ago