GetChain News
中简 中繁 EN
GetChain News
Toggle sidebar

Online/Update

News linked to both this project and an event.

Anthropic Upgrades Claude with Budget Controls, Regional Inference, Skill Loading, and Model Advisor Features

Odaily News – Anthropic has upgraded Claude this week, introducing four significant updates to Claude Managed Agents to further enhance the controllability, deployment flexibility, and task execution capabilities of enterprise-grade AI agents, including:1. New session budget control feature: Users can now set budget limits for agent sessions to achieve more precise cost management. When a session reaches its budget limit, the system will trigger an event and pause operations; users can increase the budget to resume task execution.2. Regional inference control capability: Users can now select where Claude Managed Agents run, with scheduling across globally available resources billed at standard rates. If restricted to run within US regions to meet regional deployment requirements, fees are 1.1x the standard price.3. Support for automatically loading Skills modules from user repositories;4. New model advisor feature: Users can configure a more powerful model as an "advisor," allowing task-executing agents to call upon the advisor model for second-opinion analysis during sessions, improving the quality of complex task handling.

AMD launches AI programming platform Instinct Coder, which can reduce enterprise AI coding costs by 70%

Odaily News AMD, the semiconductor giant, announced the launch of its enterprise-grade AI programming platform, AMD Instinct Coder. The platform combines AMD chips, Supermicro servers, and Spectro Cloud software, aiming to help enterprises deploy AI coding assistants locally, reduce the cost of cloud-based AI models, and protect code and data security.AMD stated that Instinct Coder is an "out-of-the-box" end-to-end AI development platform, integrating AMD EPYC processors, AMD Instinct GPUs, Supermicro AI servers, Spectro Cloud PaletteAI Inference Launchpad software, and the AMD-optimized GLM-5.2 model. It can be used for software development scenarios such as code generation, application modernization, automated testing, and code review.AMD said that compared to relying on cutting-edge cloud-based AI models, Instinct Coder can help enterprises reduce total cost of ownership (TCO) by up to 70%, with the fastest payback period shortened to 6 months.AMD noted that more and more enterprises are looking to leverage AI to improve development efficiency, but face two major challenges: on one hand, the cost of invoking top-tier cloud models continues to rise; on the other hand, entrusting enterprise source code, intellectual property, and sensitive data to third-party services poses security and compliance risks.Through a local deployment model, Instinct Coder allows enterprises to maintain control over their data and code while providing more predictable infrastructure costs. The platform supports development tools such as Claude Code, OpenAI Codex, Visual Studio Code, and Cursor, with each node supporting up to 50 users (30 concurrent users).Additionally, the PaletteAI Inference Launchpad provided by Spectro Cloud enables AI workload management, model routing, request auditing, and cost monitoring, and supports invoking external models such as Anthropic, OpenAI, Google, or xAI when needed.AMD stated that Instinct Coder aims to help enterprises break free from the high costs of cloud-based AI services, accelerate AI-driven software development processes while ensuring data security and autonomous control.

AI Reasoning Startup General Compute Secures $400 Million Loan Backed by ASIC Chips for Inference

: AI reasoning cloud startup General Compute has obtained a $400 million loan from Upper90. This deal is the world’s first financing project to use dedicated inference chips as collateral. The company has built a proprietary AI reasoning cloud platform based on SambaNova’s self-developed ASIC chips, primarily targeting Agent-type AI computing workloads. Compared to traditional GPU clouds, it offers faster token processing speeds and lower operational latency. The hardware requires no water cooling and can be directly deployed in traditional data centers and idle cryptocurrency mining facilities.

Meta's self-developed AI chip "Iris" is planned to start mass production in September, with a 2027 compute target of 14 gigawatts.

According to Reuters, Meta plans to mass-produce its self-developed data center AI chip "Iris" starting from September, as part of its fourth-generation Meta Training and Inference Accelerators project, to enhance the AI capabilities of platforms such as Facebook and Instagram and reduce reliance on external GPUs such as those from Nvidia and AMD. Internal memos show that Iris completed testing in just 6 weeks with no major defects; Meta plans to deploy 7 gigawatts of computing power this year and increase it to 14 gigawatts by 2027, with its AI infrastructure spending in 2024 potentially reaching up to $145 billion. To secure expansion, the company has signed long-term supply agreements with Samsung Electronics, Sandisk, and Sumitomo Electric to cope with "price increases" and shortages of memory and AI chips.

DeepSeek Self-Develops AI Inference Chip, Plans to Break Dependence on NVIDIA and Huawei

According to Reuters, Chinese AI startup DeepSeek is developing its own AI chips, three informed sources revealed. The chip is designed specifically for inference scenarios, rather than for model training. The project was launched approximately one year ago and remains in the early stages. The company has engaged with chip design, wafer foundry, and storage enterprises, and has quietly increased recruitment of chip design engineers without publicly posting job listings. If successfully developed, DeepSeek will reduce its reliance on Nvidia and Huawei Ascend chips, following the trend of global AI giants such as OpenAI and Anthropic developing their own hardware. Affected by U.S. export controls, DeepSeek previously shifted from Nvidia H800 to Huawei chips. This self-developed chip is regarded as a significant strategic transformation. Meanwhile, DeepSeek also plans to complete its first round of external financing, with a fundraising scale of approximately $7 billion, and a valuation reaching $52 billion to $59 billion.

Sources: NVIDIA plans to pitch Vera AI CPU to Chinese clients, some cloud providers eyeing test deployment

sources say NVIDIA has begun pitching its first independent central processing unit (CPU) product, Vera, to Chinese clients. Designed specifically for Agentic AI systems, the chip has entered mass production, marking NVIDIA's attempt to further expand its presence in the Chinese market with a CPU offering.According to sources, some Chinese clients have already shown interest in Vera. One major Chinese cloud computing company plans to procure over 300 servers equipped with dual Vera CPUs for testing, and will decide whether to expand procurement after the tests are completed.Built on the Arm Holdings architecture, Vera is NVIDIA's first independent CPU product. NVIDIA has previously stated that Vera's performance in AI agent-related computing tasks is 1.8 times that of comparable competitor products, and expects the product to contribute approximately $20 billion in revenue by the end of this fiscal year (ending January next year).The report notes that as the AI industry's focus gradually shifts from model training to inference computing, CPUs and custom chips are gaining more attention. Vera also positions NVIDIA to directly compete with Intel and Advanced Micro Devices (AMD), which have long dominated the server CPU market.Sources indicate that due to strict U.S. export restrictions on high-end GPUs, CPUs face relatively smaller regulatory hurdles in the Chinese market compared to GPU products. Currently, some Chinese clients plan to first deploy Vera chips for testing in overseas data centers. Meanwhile, software ecosystem compatibility and existing domestic AI chip deployment frameworks may still impact the subsequent large-scale adoption of Vera. (Reuters)

AMD Announces $2.5 Billion Investment in UK AI Infrastructure, Collaborating with Startup Oriole to Deploy the World’s First All-Photonic-Network AI System

According to Tech Funding News, AMD CEO Lisa Su announced at London Tech Week that the company will invest up to £2 billion in UK AI infrastructure over the next five years, covering national supercomputing infrastructure development and university research collaborations. Meanwhile, AMD is partnering with Oriole Networks—a startup spun out from University College London (UCL)—to deploy the world’s first large-scale, all-photonic network AI system under the UK government’s £50 million ARIA Inference Scaling Lab initiative. This system integrates Oriole’s PRISM photonic networking platform with AMD Instinct GPUs and EPYC CPUs; by completely eliminating electronic switches from the network core, it reduces core network energy consumption by 81% and cuts GPU idle time from 60% to under 1%.

NVIDIA Launches Nemotron 3 Nano Omni Model, Boosting Multimodal Inference Efficiency by 9x

NVIDIA announced on X platform that it has launched the open-source multimodal model Nemotron 3 Nano Omni today. The model adopts a 30B-A3B mixture-of-experts (MoE) architecture, supports a 256K context window, and can uniformly process video, audio, image, and text inputs. Compared to open-source omnimodal models at a similar interaction level, this model achieves up to a 9x increase in throughput, significantly reducing inference costs and improving scalability. Nemotron 3 Nano Omni is now available on Hugging Face, OpenRouter, and NVIDIA NIM, and has been adopted by enterprises including Aible, Applied Scientific Intelligence, and H Company.

SK Telecom Collaborates with Arm and Rebellions to Jointly Develop AI Data Center Inference Solutions

According to Yonhap News Agency, SK Telecom announced the signing of a trilateral memorandum of understanding (MOU) with UK-based chip design company Arm and Korean AI chip startup Rebellions to jointly develop AI data center inference server solutions. Under the agreement, the three parties will integrate Arm’s newly launched AGI CPU with Rebellions’ AI acceleration chip—RebelCard, scheduled for launch in Q3 this year—to jointly develop AI inference servers, which will be tested and validated at SK Telecom’s AI data centers. The Arm AGI CPU is optimized for high-density inference environments and large-scale AI deployments, while the RebelCard is specifically designed for large-scale AI inference.