Jalapeño’s first results show industry-leading speed and efficiency in AI inference ↗
OpenAI published the first performance results for Jalapeño, its custom inference chip. On GPT-OSS 120B, it says the chip delivered roughly 1.9x more peak throughput per kilowatt and around 1.7x lower end-to-end latency than the comparison system.
The gains carried across DeepSeek R1 and Kimi K2.5 as well. OpenAI plans to deploy Jalapeño inside its own compute infrastructure, while second- and third-generation designs are already underway... so yes, OpenAI is increasingly becoming a chip company too, sort of. (OpenAI)
Claude Cowork finally remembers what you told the app in chat ↗
Anthropic unified Claude's memory across regular chat and Cowork. Information learned during conversations can now follow users into agentic tasks, removing that mildly maddening ritual of explaining the same project again and again.
Users can inspect, edit and delete remembered topics, while memories update as conversations are still happening. The feature is enabled by default across Free, Pro and Max plans. (TechCrunch)
Google Cloud Launches Gemini Enterprise for Legal ↗
Google Cloud unveiled a purpose-built Gemini Enterprise package for lawyers, pushing its agents beyond generic chat and into contract review, legal research, regulatory monitoring, privacy requests and document drafting.
It includes connectors for platforms such as Harvey, iManage, Docusign, RelativityOne and Thomson Reuters, while preserving firms' existing access controls and ethical walls. The AI-agent land grab has reached the legal department... predictably, but still pretty fast. (Google Cloud Press Corner)
Google Cloud Launches Gemini Enterprise for Financial Services ↗
Google also launched a finance-specific Gemini Enterprise offering, including a managed Financial Research agent, more than 50 specialised skills and 13 connectors to financial and market-data systems.
The platform targets workflows including credit-risk assessment, portfolio monitoring, KYC research and market analysis. CME Group and Deutsche Bank are among the organisations already involved - enterprise agents are becoming increasingly specialised, almost like software wearing a suit. (Google Cloud Press Corner)
NVIDIA Announces Jetson Orin Nano 2 Robotics Computer to Redefine Entry-Level Edge AI ↗
Nvidia introduced Jetson Orin Nano 2, a compact computer aimed at robots, drones and other physical-AI systems. It offers 78 trillion operations per second, 8GB of memory and an eight-core Arm CPU.
Nvidia says inference performance doubles versus the previous Nano Super, while matching its performance can require 40% less power. Smaller models becoming genuinely capable on compact hardware is the interesting bit here - the AI doesn't always need to phone the giant data centre anymore. (NVIDIA Newsroom)
Perplexity partners with Nvidia to launch Portable Computer, a fully local AI agent with zero token costs ↗
Perplexity launched Portable Computer, bringing its agentic Computer experience onto Nvidia-powered local hardware. Models, files and agent execution can stay on the user's machine, with locally completed work consuming no cloud billing credits.
Tasks begin locally, and the system asks permission before escalating a step to a frontier cloud model. It launches around Nvidia DGX Spark and Linux systems with sufficiently powerful RTX GPUs - privacy and cost suddenly become part of the local-AI sales pitch, not just speed. (Venturebeat)
Anthropic expected to tell investors it sees over $30 trillion in potential revenue, WSJ reports ↗
Anthropic is reportedly preparing to tell investors that its total addressable market could exceed $30 trillion. Important distinction - that's the theoretical annual opportunity if AI captured the relevant work, not Anthropic predicting it will go on to make $30 trillion. (Reuters)
The enormous figure helps frame Anthropic's infrastructure spending, its competition with OpenAI and Google, and the broader investor story. It's a spectacularly large number... even by AI-industry-number standards. (Reuters)
FAQ
What do OpenAI’s Jalapeño chip results mean for AI inference?
OpenAI says Jalapeño delivered higher peak throughput per kilowatt and lower end-to-end latency than its comparison system while running GPT-OSS 120B. Similar gains were reported with DeepSeek R1 and Kimi K2.5. The results suggest OpenAI is increasingly designing hardware around the particular efficiency and speed demands of large-scale AI inference.
How does Claude’s shared memory work across chat and Cowork?
Claude can now carry information learned in regular conversations into Cowork tasks, reducing the need to explain the same project or context repeatedly. Users can inspect, edit and delete remembered topics, while those memories can continue updating as conversations develop. According to the announcement covered in the article, the feature is enabled by default across Free, Pro and Max plans.
What can Gemini Enterprise for Legal and Financial Services do?
Google Cloud’s specialised Gemini Enterprise offerings are aimed at industry-specific workflows rather than generic AI chat. The legal package covers areas such as contract review, research, regulatory monitoring, privacy requests and drafting, while the financial-services version supports workflows including credit-risk assessment, portfolio monitoring, KYC research and market analysis. Both can also connect with existing professional data and software systems.
Why is local AI inference becoming more important?
Local AI inference can strengthen privacy, reduce dependence on cloud infrastructure and potentially lower ongoing usage costs. Nvidia’s Jetson Orin Nano 2 targets robots, drones and other edge systems, while Perplexity’s Portable Computer keeps models, files and agent execution on compatible local hardware. In many workflows, cloud models can then be reserved for tasks that demand additional capability.
How does Perplexity Portable Computer decide when to use cloud AI?
Tasks are designed to begin on the user’s local Nvidia-powered hardware, keeping compatible model execution and files on the machine. If a step requires a frontier cloud model, the system asks the user for permission before escalating it. Work completed locally does not consume cloud billing credits, making the approach relevant to users concerned with both privacy and recurring inference costs.
What does Anthropic’s reported $30 trillion AI market estimate signify?
The reported figure refers to Anthropic’s view of the potential total addressable market for AI, not a prediction that Anthropic itself will generate $30 trillion in annual revenue. It represents the theoretical value of work that AI could potentially address across the economy. The estimate also helps explain why leading AI companies are investing so heavily in infrastructure and enterprise capabilities.