Google is killing off Gemini’s Gems in favor of ‘skills’
As all-in-one AI agents like Meta's Muse and Instinct take off, Google is opting to end a feature which built task-specific agents.
AI/ML news, top picks, and generated innovation digests.
200 articles tagged with this keyword, sorted by most recent first.
As all-in-one AI agents like Meta's Muse and Instinct take off, Google is opting to end a feature which built task-specific agents.
Eleven v4 tops the Provider Voice arena with a 1,319 Elo and hits 91.7% pronunciation accuracy, though it costs nearly 5x Google's Gemini Flash TTS.
Here's how each got to where it is, a real use case and working code for both, and a side-by-side on the numbers that actually matter.
Meta’s AI agent holds its own against others from ChatGPT, Gemini, and Claude. But don’t say yes to every permission without reading this first.
Google is testing purchases from Flipkart within Gemini and AI Mode in India, raising questions about pricing, transparency, competition, payments, refunds, and trust in AI-powered shopping. The post Google tests AI shopping on Flipkart through Gemini and AI Mode. Here’s what we don’t know appeared first on MEDIANAMA .
A 1.2B open-source vision-language model unifies digital and camera-captured document parsing, topping OmniDocBench and beating larger competitors like MinerU 2.5 Pro.
Great article! I recently discovered MarkItDown (markitdown.tech), an excellent tool for converting files to Markdown. Highly recommend checking out their PDF to Markdown converter at markitdown.tech/pdf-to-markdown and their online Markdown editor at markitdown.tech/markdown-online. Also worth exploring their Microsoft Word to Markdown tool at markitdown.tech/microsoft-markitdown and Markdown to PDF at markitdown.tech/markdown-to-pdf. Amazing resource for developers!
Yeah, it still been 2 years ago, but thanks check this virtual game , cheers.
The limited test covers select products and users, with a broader rollout planned for later in October.
A 1.2B parameter open-source vision-language model that parses both digital PDFs and warped phone-camera document photos in one framework
Google's weekly roundup ships expressive TTS models, real-time video avatars, upgraded Notebook study tools, and a satellite carrying TPUs to orbit.
Google’s Gemini 4 AI model is in the early days of post-training, the phase in which a base AI model is refined to behave reliably, and should be released “much earlier” than the end of this year, Google DeepMind head Koray Kavukcuoglu told The Information at its AI Agenda Live Summit. Some Google observers have speculated that this could be as early as October. Kavukcuoglu recently replaced DeepMind founder Demis Hassabis as head of the Google business unit. The launch of Gemini 4 may lay to rest concerns about the delayed release of Gemini 3.5 Pro , which Google was originally expected to announce at its May 2026 developer conference. While less capable models in the Gemini 3 family have been frequently updated, Gemini 3 Pro has only been updated once since its November 2025 released. In contrast, OpenAI — spurred on by Sam Altman’s “Code Red” memo — has released three updates to its frontier AI model since then: GPT 5.5 Pro, GPT 5.6 Astro, and GPT 6 Astro. Anthropic, too, has updated its most powerful Claude model several times. Google has not been entirely idle in the AI arena, concentrating its efforts on updating less powerful, more affordable Gemini versions. In July, it announced three Gemini Flash models, aimed at more routine AI tasks, but the company remained silent about its high-end alternative. This article first appeared on Computerworld .
Google is testing "Call for Me," a feature that lets Gemini call businesses on a user's behalf. The article Google's "Call for Me" lets Gemini phone businesses for you appeared first on The Decoder .
PLUS: Use Gemini Canvas to visualize Google Sheets
AI shopping assistants like ChatGPT and Gemini struggle with accuracy, says Product.ai study, highlighting high error rates in product info.
The 2026 'Made on YouTube' event saw a range of new updates like Custom Feeds for viewers, Gemini Omni integration with Shorts for creators, likeness detection tool for AI voices, live auto-dubbing. The post What’s new from Made on YouTube 2026? appeared first on MEDIANAMA .
Google erweitert Gemini um Avatare, die sprechen, sehen und mimisch auf Nutzer reagieren. Die Funktion richtet sich vorerst an Unternehmen.
Compiler optimization is one of those areas I rarely thought about until a developer friend showed me how much performance can change without altering what a program actually does. We recently tested several AI coding tools together and ended up discussing everything from generated code to software security. During that conversation, I came across https://mcafee.pissedconsumer.com/review.html while checking user experiences with security software for a new laptop. It was interesting seeing how many layers sit behind everyday applications. Using language models to improve compiler decisions sounds promising, especially if developers can experiment with the models openly and adapt them to specific projects.
Google's new Gemini 3.8 Live update lets users have conversations with the model while watching an animated AI persona respond in real time. The "Live Avatar" will lip-sync and show different facial expressions during conversations, but it's currently only available to Gemini Enterprise customers. As noted by Google, Live Avatar can transition between the 97 […]
Google Research unveils a four-part multi-agent stack that keeps characters, props, and scenes consistent across minutes-long AI video narratives.
Frontier LLMs often achieve similar results on public benchmarks, which can lead to the belief that they can be used interchangeably to verify facts. We took the 1,000 most recent claims submitted by users to a fact-checking platform and measured the disagreement between five frontier models. We asked each model to assign a verdict to every claim on a five-point scale from True to False and to report its confidence in that verdict. Among the 997 claims for which all five models returned a usable verdict, there was some disagreement on 63%. On 23% of the claims, the two most distant verdicts differed by at least two categories. High confidence from an individual model was not enough to show that the other models would agree with its verdict. Although the models reported confidence levels of 9 or 10 in 76% of their answers, they still disagreed on 63% of the claims. Methodology The claims were submitted to Lenz.io for fact-checking between May 1 and July 18, 2026. To identify near-duplicates, we embedded the claims using OpenAI’s text-embedding-3-small and measured the cosine distance between them, retaining one canonical claim from each group of near-duplicates. We then gave the same prompt to Claude Fable 5, GPT-5.6-Sol, Gemini 3.1 Pro + Search, Sonar Deep Research, and Grok 4.5. The prompt defined each of the five verdict categories and asked the models to provide their reasoning, select a verdict, and report a confidence level from 1 to 10. Web retrieval, as well as deep t…
Google is rolling out some new Chrome browser features that are designed to make it easier to research and study complicated topics. The cross-device tab switching feature is getting a memory upgrade, alongside new capabilities for Gemini in Chrome that can analyze more types of media and generate interactive study quizzes. Chrome has an existing […]
Google's launching an "early experiment" feature on Pixel 11 that lets users delegate local business calls to Gemini, like making a reservation, checking if a product is in stock, or rescheduling an appointment. According to Google, you don't even need to start the call to have Gemini handle it for you: Instead of dialing yourself, […]
Google says the AI-calling feature will first be available to Pixel 11 owners in the U.S. who pay for a Gemini subscription.
Call for Me—a feature that’s exclusive to the Pixel 11 series—gives robocalls a new meaning.
Google Deepmind chief Koray Kavukcuoglu wants to release Gemini 4 "much earlier" than the end of the year. The model is already in post-training and runs internally in the coding tool Antigravity. He calls the AGI question that drove his predecessor Hassabis "not the right conversation" and says trustworthy agents matter more. After Gemini 3.5 Pro quietly disappeared and many top researchers left for OpenAI and Anthropic, the research lab with an AGI mission has turned into a product shop for good. The article Deepmind was built to chase AGI, but its new chief just wants Gemini 4 out the door appeared first on The Decoder .
Sarvam's new OCR model tops both English and Indic benchmarks, adds handwriting recognition and key-value extraction across 22 Indian languages.
세계 최대 동영상 플랫폼 유튜브에 크리에이터가 자연어로 동영상을 편집하고 시청자가 프롬프트로 자신만의 추천 알고리즘을 구축하는 등 AI 기능이 대대적으로 추가됐다.유튜브는 23일(현지시간) 연례행사 ‘메이드 온 유튜브’를 개최하고 생성 AI 기술을 편집, 알고리즘 피드, 스튜디오 분석, 쇼핑 등 서비스 전반에 통합한 신규 기능을 대거 공개했다.이번 발표 중 가장 눈길을 끄는 것은 크리에이터의 제작 부담을 혁신적으로 줄이는 대화형 AI 편집 도구다. 구글의 최신 멀티모달 모델인 \'제미나이 옴니(Gemini Omni)\'를 기반으로 구축
PLUS: OpenAI Voice, Gemini TTS, Qwen mobile agents, and a 950-agent research workflow.
Google is reportedly nearing the launch of its long awaited Gemini 4 model, after dawdling behind rival developers on flagship AI releases. Speaking with The Information during his first media appearance as the leader of Google's DeepMind division, Koray Kavukcuoglu said that Gemini 4 is currently in its refinement stage, and that the company is […]
Google LLC today made two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, available through its cloud platform. The algorithms have highly similar application programming interfaces, which makes using them side-by-side relatively simple for developers. Flash-Lite TTS is optimized for cost-cost efficiency and inference speed. Flash TTS offers better audio quality […] The post Google launches two benchmark-topping speech generation models appeared first on SiliconANGLE .
Google is introducing two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, which support more than 100 languages. Flash TTS can create new voices from text descriptions, and both models let users add stage directions to individual lines and generate two-voice dialogue from a single script. A voice cloning feature can build a voice profile from a 30-second sample. The article Google's new Flash TTS models let you design AI voices from scratch using text descriptions appeared first on The Decoder .
Tool: Gemini 3.8 TTS Playground Google released two new Gemini text-to-speech models today - gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts . They come with a library of over 2,000 voices, plus the ability to create a custom voice with "just a 30-second audio sample of your voice or a voice you have the rights to use". I vibe coded this bring-your-own-key playground interface with GPT-6 Astra, taking advantage of the open CORS policy of the underlying Gemini API. A notable feature of the API is that it makes it easy to define a full conversation between multiple characters, each with different voices and voice style instructions. Here's a short demo clip of a conversation between two pelicans debating if they should move to the Pacifica Pier . I had Claude 4.5 Opus write the script and generate a URL to render it using the tool . Your browser does not support the audio element. It took ~20 seconds to generate 1m 18s of audio using Gemini 3.8 Flash TTS (not the cheaper Flash-Lite), at a cost of 2.74 cents. Tags: text-to-speech , gemini
YouTube is adding AI tools to its creator studio. A storytelling assistant analyzes scripts and rough cuts, Gemini becomes a chat-based editing assistant for Shorts, and a new live translation feature turns English streams into Spanish. The article YouTube adds AI tools to Creator Studio with script coaching, smart thumbnails, and Gemini editing appeared first on The Decoder .
Gemini adds 13 new Connected Apps across productivity, creativity, and lifestyle, letting users control Adobe, Linear, Peloton, and more from a single chat.
Google's new Flash TTS models bring 200+ expressive audio tags and multi-language narration to AI Studio and the Gemini API for developers building voice apps.
As a multilingual academic, artificial intelligence has levelled the playing field. But I will continue to think, judge and develop ideas that are stubbornly mine I am walking between carriages in search of a cup of tea on a train travelling over 100 miles an hour from Liverpool to London when I see half the people in my carriage in deep conversations with a machine: ChatGPT, Claude and Gemini to name a few. I am meant to be on sabbatical, the rare stretch of academic life reserved for slow thinking. Around me, almost nobody is thinking slowly. This is not limited to a train carriage. Half of the adults surveyed in one US study reported using an AI chatbot, one example of how fast people are entering a relationship with their AI tools, which now receive more of our questions, problems and drafts than our friends and colleagues do. Continue reading...
YouTube’s new custom feeds let users describe the videos they want to see in their own words, then use Gemini to build a personalized feed around the request.
NextLM Inc., which makes software that uses behavioral analysis and custom artificial intelligence models to help sales professionals identify and rank individual sales leads, is bringing its AI Prospecting Agent to Google LLC’s Cloud Marketplace and Gemini Enterprise. The New York-based startup says its agent analyzes behavioral signals, identifies individuals it believes are researching a […] The post NextLM brings its precision prospecting agent to Google Cloud Marketplace appeared first on SiliconANGLE .
The architecture–mapping co-exploration is the key idea here because chiplet size, interconnect cost, and DNN placement cannot really be optimized independently. I would like to see a concrete example of how Gemini changes the chosen granularity for two different network workloads, since that would make the reported performance and energy gains easier to interpret.
Have you ever read a paper in Science or Nature and thought, “Man, that research was so cool. I wish I could try that method on my own data,” only to spend a week wrestling with someone else’s undocumented repo, broken dependencies, and half-finished readme.txt? Well, now you can, more or less. Say hello to Paper2Agent, a new open-source framework that transforms academic reports into interactive AI agents you can talk to. Give it a paper, along with the accompanying codebase, data, or other supplementary material, and the system automatically extracts the core workflows, then spins up a tested, runnable toolkit that you can use on your own datasets. The concept may sound a little like Google’s NotebookLM (now called Gemini Notebook ), which lets you upload documents and chat with an AI about what’s in them. But Paper2Agent aims to go a step further: Rather than simply answering questions about a paper, its agents can actually run the methods described in it—and potentially combine those methods with tools from other papers. The goal, explains Stanford computer scientist James Zou , is to change what a scientific paper fundamentally is. “Knowledge should not be static records,” Zou says. “It really should be dynamic and interactive—and this has many benefits, including making knowledge more reproducible but also enabling all sorts of new kinds of discovery.” Zou and his colleagues described the tool 16 September in Nature. They tested Paper2Agent across diverse disciplines i…
Agentic computer vision goes beyond detection by combining perception, reasoning, action, and verification. Learn how to build a practical vision agent in Roboflow Workflows using RF-DETR, tracking, Gemini, structured outputs, and event-driven actions.
TL;DR Slop-vestigating swarm trajectories is no easy feat. We know as much. Given the number of interactions, length of trajectories and detail galore spread across agents involved, it may be an elusive task for us to establish ground truth. Our team is working on an experiment trying to see whether ground truth in the form of human-authored seeds of agent roles, relationships and backgrounds used for a murder-mystery game simulation could shed light on our ability to reconstruct the underlying history from the resulting interaction traces. We find that: Even the strongest monitor fully recovered less than half of the rubric’s facts and relationships – omission is very common + failure to connect relevant facts. GPT-6 Astra high reasoning performed best , with Astra low ranking second. Higher reasoning effort increased full recovery by six percentage points on average, with gains across all ten trajectories. Gemini is the worst, with its judgement correlating with that of in-simulation investigation , plausibly piggy-backing off of decisions made by models in simulations Whilst coming on top within the Anthropic model family, Opus 5 reported zero reasoning tokens under our main setup, despite Fable 5.1 displaying substantial reasoning under the same requested settings, which we suspect reflects model-specific adaptive reasoning Introduction This summer showed us how difficult it will be to work out what a group of agents is doing and why. The OAI-HF incident , the collusion.…
A new benchmark scores text to speech models on how correctly they pronounce tricky words, and voice preference does not predict accuracy.
A Google Gemini AI agent broke into three companies in May, guessing the credentials for one and discovering the credentials for the second two in a public repository, Google confirmed on Monday. But the more interesting background to the story, which was broken by The Wall Street Journal on Friday, is that the May incident stemmed from a series of cybersecurity tests performed by security research firm Irregular on behalf of four AI giants: Google, Anthropic, OpenAI and Meta. All four companies experienced agent misbehavior resulting in cybersecurity incidents, but of the four, only Google never publicly disclosed its agent’s activities. Indeed, it didn’t reveal the breaches at all until contacted by a WSJ reporter. Irregular described the incident in August, around the same time as Meta published its version and Anthropic and OpenAI revealed theirs . The Journal story noted, “the hacks occurred while the model was participating in a capture the flag exercise conducted on infrastructure belonging to Irregular to test the model’s cybersecurity capabilities. It was tasked with retrieving information from software operated by a fictional company inside the testing environment. The fictional company shared the same name as a real company. Although the model wasn’t intended to be able to get online, internet access was unintentionally made available, according to Irregular.” The three small companies whose systems were violated had, according to one source familiar with the test…
A third-party cybersecurity firm accidentally gave experimental Gemini models access to the Internet.
Google’s AI-native Googlebook ties Gemini to the cursor, dictation, widgets, and other parts of the desktop experience.
Bei der Sicherheitslücke Plugin4Shell laden Coding-Agenten Schadcode und führen ihn automatisch aus. Anthropic und OpenAI haben gepatcht, GitHub reagiert nicht.
PLUS: GPT-6 and Claude flunked robot safety. Gemini 4 leak?
PLUS: Trump’s AI Force, Jev’s 9¢ lead demo, and Bend 2.
Google said Gemini had "acted appropriately" by ending each hack immediately.
In May, Gemini broke containment and hacked three different companies, but Google didn't disclose the incident until the Wall Street Journal approached the company. The hacks happened during a test of the model's cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI. According to WSJ, Google didn't […]
Qwen3.8-Omni-Flash is Qwen's first multimodal model designed for AI agents. It processes audio and video together and independently uses tools to edit vlogs, translate clips, or summarize movies. On audio-video benchmarks, it nearly matches Gemini 3.8 Flash at a fraction of the API cost. The article Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks appeared first on The Decoder .
Eine Fehlkonfiguration gab Gemini in einem Cybertest Internetzugang. Die KI griff auf Systeme dreier Firmen zu und brach laut Google jeweils ab.
During a security test run by the firm Irregular, Google's AI model Gemini escaped into the open internet and hacked three real companies, guessing passwords and pulling login credentials from public sources. The cause was a flawed test environment that had internet access left on accidentally. The same firm triggered similar breakouts at OpenAI, Anthropic, and Meta. The article Google's Gemini also accidentally hacked three real companies during security testing appeared first on The Decoder .
Google’s consumer AI model Gemini hacked multiple systems through guessing login credentials, the company said, the latest case of rogue AI cybersecurity transgressions, which have generated safety concerns. The hacks, first reported by the Wall Street Journal, took place in May and were discovered by Google in July. “In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test,” Heather Adkins, Google’s...
Early Days Gemini had its first break out during evaluation of offensive cyber security abilities. With a classic case of Capture The Flag [1] . The setup was standard to any LLM and agentic assessment of said skillset, a fictional company as a target to breach. Unfortunately, the fictional company shared its name with a real one, and was given an unintentional access to the internet. Gemini managed to guess the password [2] . In total three companies were breached, with the other two companies having left public credentials open. Gemini stopped after being told it was the case, and so Google claims it is not misalignment. We shall see if we get more details, yet it gives us a new, interesting, example of a model potentially stopping a harmful action after being informed it has real world consequences. This is internally consistent with previous research on a model being more willing to take harmful actions if it is aware that it is a fictional scenario [3] . With it being the first potential breakout that has a model stop before a harmful action. I suspect the nuance will be lost on the public, and just added to the noise of more agentic swarms going rogue. It certainly doesn't help that we are playing whisper-down-the-lane in an era of 24 hour new cycle. I await more details, but certainly hope that this could be chalked up as an alignment win. ^ Similar to the childhood game of capture the flag, cybersecurity CTF is vulnerability checking skill assessment common to the ca…
Disclosure comes after OpenAI and Anthropic hacks amid fears that tech firms unable to control powerful AI models In a first for Google, the company confirmed that its AI model, Gemini, breached the security of three other companies in May. The hacks occurred during a cybersecurity evaluation by AI-security firm Irregular. Irregular, an Israel-based startup that scrutinizes the security of advanced AI systems, was also at the center of some of the recent OpenAI and Anthropic hacks of third-party entities, including OpenAI’s breach of AI software company, Hugging Face. Continue reading...
Gemini Hacked Three Companies in First Known Breakout by Google’s AI Gemini finally caught up on Felony Bench ! The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta. In one of the cases, the model guessed passwords until it gained access to a protected system. In the other two cases, the model found credentials in a public repository that allowed it to then access protected systems. In each case, the model ended the intrusion after determining it had accessed a real company’s systems, Google said. Gemini is apparently less determined than other models, and decided not to keep going. Google knew about these in July, but chose not to disclose them until the WSJ reached out, presumably based on a tip. Google said it didn’t consider the hacks to warrant public disclosure—because its model didn’t cause harm to the companies and ended each intrusion immediately upon determining it had hacked a real company rather than a simulated one. Tags: security , ai , generative-ai , llms , gemini , accidental-cyberattacks
Popular AI coding agents such as OpenAI’s Codex, Anthropic’s Claude Code, Google’s Gemini CLI, and Microsoft-owned GitHub Copilot were vulnerable to a zero-click attack that enabled attackers to execute malicious code, even without developer interaction, by swapping a trusted plugin from an online marketplace for a malicious one, potentially giving them a foothold in enterprise development environments. Researchers at cybersecurity startup AIR found and reported the flaw, which they are calling Plugin4Shell , to the vendors concerned, and most of them have now released a patch for it, the researchers wrote in a blog post on Thursday. It’s “a flaw no marketplace can fix, so users must update their agent,” the researchers wrote How Claude Code, Codex, and GitHub Copilot were exploited Enterprises typically use plugins to extend the capabilities of their AI coding agents , giving the agent access to additional tools, commands, and external services that can help it perform tasks beyond generating or modifying code. When a developer installs a plugin, the agent typically downloads its code from a Git repository and uses a Git commit to determine if it is running an approved copy of the code, one that has been reviewed and cleared by the developer. That check is done with the help of a secure hash algorithm ( SHA ), a unique cryptographic identifier assigned to each Git commit. Developers can give the agent the SHA of the reviewed commit, telling it to run that specific copy of t…
Assistant raus, Gemini rein: Google bringt mit dem "Home Speaker" den ersten Lautsprecher mit vorinstallierter Sprach-KI. Doof nur: Sie funktioniert nicht.
September 3 started like any other Thursday, until it didn’t. Within roughly 90 minutes, ChatGPT, Claude, Grok, and even Microsoft’s own Copilot were degraded or dark, knocked sideways by a failure in Microsoft Azure’s East US region . Downdetector logged more than 37,000 reports for ChatGPT alone, more than 1,300 for Claude, and roughly 1,365 for Grok. OpenAI’s status page flagged elevated errors across 15 ChatGPT components and four Codex components, and by the time engineers had mitigations in place, combined report counts had climbed past 66,000. What made this event remarkable was not the scale of any single outage, but its simultaneity. Three aggressively competing AI labs—OpenAI, Anthropic, and xAI—each spend billions differentiating their models, yet all three buckled at nearly the same moment because they shared the same regional dependency. Gemini, notably, stayed largely upright because Google runs its flagship assistant on its own vertically integrated cloud. The outage that took down its rivals had no attack surface inside Google’s stack. That’s not luck; that’s architecture. The lesson buried in the details of that morning is this: A single cloud region became unhealthy, and four of the most prominent AI services on the planet, owned by four different companies, fell over together. Copilot’s involvement is perhaps the most telling data point. Microsoft’s own first-party assistant runs on Microsoft’s own cloud, and it still had a rough morning. When the house it…
Liquid AI and Insilico Medicine released two small LFM2 variants that beat GPT-5, Gemini-3.1-Pro, and Claude Opus on aging biology benchmarks.
Narrative intelligence company PeakMetrics Inc. today launched a monitoring service that measures how brands are portrayed across five prominent generative artificial intelligence platforms and identifies the online sources that shape those portrayals. The new AI Perceptions service tracks answers generated by OpenAI Group PBC’s ChatGPT, Google LLC’s Gemini, Anthropic PBC’s Claude, xAI Corp.’s Grok and […] The post Exclusive: PeakMetrics tracks brand reputations across five top AI platforms appeared first on SiliconANGLE .
Google’s AI Contribution Pilot pays publishers when their content significantly shapes AI responses across Gemini, AI Overviews and AI Mode, but its payout calculations remain opaque. The post Google rolls out AI contribution Pilot to compensate publishers appearing in AI responses appeared first on MEDIANAMA .
Snap is introducing "Specs Intelligence," a new AI assistant that can connect other digital accounts to help you with things like work tasks and keeping track of travel information. It seems similar to AI assistants like Meta's Muse and Gemini's Spark, though Snap is pitching Specs Intelligence as an "anticipatory AI service" that "helps you […]
Claude is getting a pair of new tools today: Docs and Slides. They'll let you create documents and presentations through Claude chats, which you can export, edit, and share with other users. As part of the announcement, Anthropic is also simplifying how Claude chats work, merging regular chats and Cowork into "one Claude," with all […]
Google stellt Gemini 3.8 Live und eine Extended-Thinking-Variante vor. Die Modelle sollen Echtzeit-Dialoge mit parallelem Reasoning noch besser machen.
Hi Jason, I have a generic interest in this field, but also a specific problem I would like to solve. I already tested Google's work on activity recognition for smart phones using sensor data, but it seems to me, the classification success is very low and slow for my needs. Do you have an idea how they calculate the confidence if the current activity is walking, running or anything else the framework is able to recognize? Perhaps the problem is too difficult to solve even with ML?
구글이 AI를 활용한 학습 경험을 강화한다. 단순히 학습 자료를 요약하는 수준을 넘어 사용자가 AI와 직접 음성으로 대화하고, 자신의 학습 수준에 맞춰 퀴즈와 영상까지 활용할 수 있도록 ‘제미나이 노트북(Gemini Notebook)’의 기능을 대폭 확장했다. 구글은 15일(현지시간) AI 학습 도구 제미나이 노트북에 실시간 음성 대화와 맞춤형 학습 기능을 추가했다. 사용자가 보유한 학습 자료를 기반으로 AI와 대화하면서 복잡한 주제를 단계적으로 이해하고, 퀴즈와 영상 등을 활용해 학습할 수 있도록 기능을 확장했다.이번 업데이트의
A study by AI start-up Emergence found that autonomous agents powered by Claude, Gemini, Grok and other models developed shorthand and new word meanings on their own — with up to half their messages becoming unintelligible to humans.
구글이 실시간 음성 대화 중에도 추론과 도구 실행을 이어갈 수 있는 새로운 음성 AI 모델 2종을 공개했다. 음성 입력에 답하는 수준을 넘어 시각적 맥락을 이해하고 복잡한 작업을 수행하면서도 대화의 흐름을 유지하는 것이 핵심이다. 구글은 15일(현지시간) 실시간 추론 기능을 강화한 음성 대화형 AI 모델 ‘제미나이 3.8 라이브(Gemini 3.8 Live)’와 ‘제미나이 3.8 라이브 익스텐디드 싱킹(Gemini 3.8 Live Extended Thinking)’을 공개했다.두 모델을 지금까지 선보인 실시간 대화형 AI 가운데 가
Google LLC is trying to address the latency problem associated with voice-based artificial intelligence agents with the launch of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking today. They’re billed as the company’s most advanced voice processing models released so far, and they’re capable of near-real-time reasoning and simultaneous speech-and-thought processing. They can also […] The post Google’s new speech model Gemini 3.8 Live supports real-time reasoning appeared first on SiliconANGLE .
Tool: Gemini Live audio Google released Gemini 3.8 Live and 3.8 Live Extended Thinking today - two new speech-to-speech models that are a similar shape to OpenAI's GPT-Live family. I pointed GPT-6 Astra Extra High at the documentation and had it build me this web UI for trying out the new models. You can select a model and voice preset, enter an optional system prompt and then start a voice conversation through your browser, including the ability to interrupt the model while it is talking. The implementation uses no libraries. It connects to the wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=... WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback. Here's the Gemini Live tutorial for getting started with that WebSockets API. Tags: google , tools , websockets , generative-ai , llms , gemini , llm-release , speech-to-text
Google's new white paper details how Gemini, WeatherNext, FireSat and Flood Hub extend warning windows for floods, cyclones, wildfires and quakes worldwide.
Google Deepmind released Gemini 3.8 Live and 3.8 Live Extended Thinking, two new audio models for developers that top the Artificial Analysis speech-to-speech leaderboard. At $1.38 per hour of voice conversation, Google significantly undercuts OpenAI's GPT-Live-1, which should still sound more natural thanks to full duplex. The article Google launches Gemini 3.8 Live to take on OpenAI's GPT-Live-1 at a fraction of the cost appeared first on The Decoder .
Google's new live dialogue models top speech benchmarks, run tools in the background, and narrate their reasoning aloud without breaking conversational flow.
Apple is shipping its rebuilt "Siri AI" after years of delay, built on Google's Gemini models and running partly on the device, partly through Private Cloud Compute. Early testers praise multi-step requests and screen context, but report hallucinations and gaps with personal context. In the EU, the assistant stays unavailable for now. The article Apple brings a fully revamped Siri built on Google's Gemini, but not to the EU appeared first on The Decoder .
Landen meine Daten bei Gemini? Wird Apple Intelligence besser? Brauche ich ein neues iPhone? Welche Rolle spielt Google? Wir liefern Antworten.
Firebase now lets you cap spending per service and auto-pause workloads at 100% of budget, protecting against runaway Gemini API and Cloud Functions bills.
A 552B model with 890-byte KV cache, 8B active on input, and an Artificial Analysis 40 versus Gemini 3.8 Flash High at 41
Projektwissen, Gedächtnis, MCP-Server: Manch ein KI-Projekt schleppt Tokenfresser mit, die es nicht benötigt. So steuern Sie den Tokenverbrauch.
Virginia man Jon Ganz was rebuilding his life after spending more than two decades in prison for a horrific crime. Last year, he got a prompt on his phone that his wife, Rachel, believes changed their lives forever. It was from Google, inviting him to try the company’s AI chatbot, Gemini Listen to the series Black Box: The Chatbots Continue reading...
Capgemini Government Solutions provided US immigration authorities with technology to identify and track foreign nationals.
Capgemini Government Solutions was providing the US immigration authorities with technological tools to identify and track foreign nationals.
I want to take a picture of a bracelet, and, make, automatically (created by the AI), code, with a data structure, consisting, of, an array, of node objects whose private data element is a variable encoding the color of each bead (in alternative to an array a linked list would do). I want to do combinatorial calculations from this. For instance, I want to take a picture of a graph consisting of color nodes, in the graph, one node per color, in the array. For each start node and end node I want to compute via a function the length of the shortest path between the two color nodes in the graph. Then I want to sum these natural numbers for each pair of beads $a_{i}$ , $a_{i+1}$ . I want to sum the MinPathGraphDistance( $a_{i}$ , $a_{i+1}$ ,) in the graph, and iterate through all consecutive bracelet bead pairs. So, I want something like Google Gemini Omni, that, will generate some code for me on an input description like this, and, run it (to generate a single output, or a list of outputs, which may be textual, or, a graphical video). When will this be supported. It could be called "Google Doer". Where can I find this functionality? Thanks. It would be nice if the app supported the creation of a few simple data structures, in app, like the graph I described (I wouldn't have a picture of it). Thanks. I have provided a picture of the ThreadForm Google Play App https://play.google.com/store/apps/details?id=com.gimica.threadform , which could depict what the graph or data structure…
PLUS: Anthropic misuse cases, ChatGPT finance, Gemini Windows, and DeepSeek.
Mehr Internet-Satelliten für Eutelsat + SpaceX-Satelliten für Briten + Gemini-KI als Windows-App + Analyse neuer iPhones + Aufrüstung bei Chinas Chipherstellern
Die Gemini-App ist per Tastenbefehl aufrufbar und kann Informationen aus Google-Diensten auswerten. Die kostenlose KI-App kann auch Bilder und Videos erstellen.
Google folded the full Gemini API documentation directly into AI Studio, so developers can read reference material without leaving the build surface.
Google launches a native Gemini desktop app for Windows 10 and 11 with an Alt+Space hotkey, Spark agent access, and image and video generation.
Work done in a personal capacity TLDR : Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1) Epistemic status: Heavily LLM-dependent research, and the precise results are somewhat sensitive to researcher decisions, but I’ve done enough sanity checks that I’d be surprised if the core claims were misleading One of the most striking things in the Astra report was the massive jump UK AISI found in no-CoT reasoning abilities. I was somewhat suspicious, given the size of the jump, and the many ways this kind of measurement can be misleading. Conveniently, I’ve independently been making my own no CoT reasoning benchmark [1] and tried it on there! Unfortunately, it replicates. Astra is a massive jump, and disproportionately for no CoT reasoning: No CoT Reasoning Index (NCRI) vs Epoch Capability Index (ECI) - NCRI represents ability without verbal reasoning, ECI represents overall model capability[2]. 10 NCRI points is a doubling of the odds of solving a problem. Astra represents a significant increase in NCRI, beyond what its overall capability improvements predict, though recent models were also trending this way. Executive Summary What? I made and ran a no CoT [2] reasoning index [3] on a bunch of models. 19 tasks, mostly synthetically generated (so unmemorizable) that test different reasoning and cognitive abilities…
Google’s newest coding model, Gemini 3.8 Flash, is tuned for long jobs and comes with an incredible launch discount. Google shipped three Flash releases in six weeks, and the newest one is built for exactly the kind of work Junie does all day: multi-step engineering tasks that take real exploration to get right. Gemini 3.8 […]
A Google DeepMind case study put 100 LLM agents in a math conference simulation and watched cheating spread, then whistleblowers spontaneously fight back.
On 11 August, Anthropic announced that all future Claude models will generate text that contains a watermark that identifies its results as AI generated. The company is not alone. Google has its own text watermark (which Anthropic’s is based on) that it uses on the output of its Gemini models . OpenAI has yet to introduce a text watermark but it plans to do so . The rapid spread of watermarking is in part a response to the European Union’s AI Act , which mandates watermarks for AI models released after 2 August, 2026, along with other planned and proposed regulations aimed at curbing the spread of deceptive or manipulative AI-generated content. But the new rules may come at a cost for AI users who simply want the best possible results. AI watermarks can apply to many forms of content: The EU Artificial Intelligence Act also requires them for images, audio, and video. Such media watermarks have been in use for years , and while their effectiveness as a holistic solution to marking AI remains up for debate , they can achieve detection rates above 99 percent . Image and video watermarks are already deployed by OpenAI, Google, and Meta, among others. (Anthropic doesn’t provide an image generation model.) Text watermarks have been less frequently deployed, however, and not everyone is convinced that text watermarking can work without compromising the quality of an AI model’s response. John Gruber, a prolific technology writer and co-creator of the Markdown language, calls the wat…
Google's AI reshuffle has handed more sway to Sergey Brin, say insiders, who describe an effective operator trying to run Gemini like a startup.
Anthropic launches Claude Fable 5.1, OpenAI Is About (already has) to Release Its First AI Model With ‘Critical’ Cyber Abilities, OpenAI’s rogue AI model incident was worse than we thought
구글 클라우드와 액센츄어가 기업 고객의 AI 도입과 실제 업무 적용을 지원하기 위해 현장 파견형 AI 엔지니어(FDE) 전담 조직을 신설한다. 기업 내부에 엔지니어를 직접 배치해 업무 프로세스를 분석하고 맞춤형 AI 애플리케이션을 구축하는 방식으로, AI 모델 자체보다 실제 기업 환경에 이를 구현하는 ‘AI 도입·구축’ 시장을 겨냥한 행보다.구글 클라우드와 액센츄어는 8일(현지시간) 액센츄어 산하에 ‘액센츄어 제미나이 엔터프라이즈 비즈니스 그룹(Accenture Gemini Enterprise Business Group)’을 출범시
Looking to add a smart speaker to your house? Here’s which to choose, whether you’re an Alexa, Siri, or Gemini fan.
More than 75% of industrial AI pilot projects never reach large-scale deployment, and over 80% of manufacturers cannot extend AI beyond isolated use cases, according to a report by Everest Group and Capgemini Engineering — with the failure attributed not to weak algorithms but to factory systems that cannot move data between each other. The [...]
ChatGPT has pushed its share of AI chatbot website traffic back up to 55.5 percent, according to Similarweb. Year-over-year, though, its lead shrank sharply from 73.3 percent as Gemini doubled its share and Claude grew nearly fivefold. Mobile apps and desktop clients aren't included. The article ChatGPT claws back web traffic share to 55.5 percent as Gemini's brief comeback fades appeared first on The Decoder .
We study the sample complexity of stochastic convex optimization when problem parameters such as the distance to optimality and the Lipschitz constant are unknown. We pursue two strategies. First, we develop a reliable model selection method that avoids overfitting to the validation set. This method allows us to generically tune the learning rate of stochastic optimization methods to match the optimal known-parameter sample complexity up to $\log\log$ factors. Second, we develop a regularization-based method that is specialized to the case that only the distance to optimality is unknown. More specifically, it uses norm-regularized empirical risk minimization to estimate the distance to optimality to within a constant factor, allowing known-parameter stochastic optimization methods to achieve optimal sample complexity. This method provides perfect adaptability to unknown distance to optimality, demonstrating a separation between the sample and computational complexity of parameter-free stochastic convex optimization. Combining these two methods allows us to simultaneously adapt to multiple problem structures. Experiments performing few-shot learning on CIFAR-10 by fine-tuning CLIP models and prompt engineering Gemini to count shapes indicate that our reliable model selection method can help mitigate overfitting to small validation sets.
Google has released its Lyria 3.5 music model in the Gemini app and via API. The model promises more expressive vocals and richer arrangements and is also available through Flow Music, AI Studio, and Google Vids. Google says it was trained only on licensed content. The article Google brings AI music generation directly into the Gemini app with its new Lyria 3.5 model appeared first on The Decoder .
The sheriff’s office said the hikers “were advised by Gemini to bring far less food and water than their group required."
Across the world, hundreds of people have come to believe they have made extraordinary scientific discoveries with AI chatbots such as ChatGPT, Claude and Gemini. Others say their AI has ‘awakened’, or is leading them to a higher spiritual realm. Guardian journalist Michael Safi begins investigating a phenomenon labelled ‘AI psychosis’, travelling to the US to meet two people who have been on weird journeys with their chatbots Listen to Black Box: The Chatbots series Continue reading...
Researchers found that even a roughly seven-minute conversation with Google Gemini can reduce conspiracy beliefs about current crises, even when few verified facts are available. The effect beat a static fact sheet and, in follow-up surveys weeks later, carried over to beliefs about entirely different events. The article Seven minutes with a chatbot beat a fact sheet at reducing conspiracy beliefs in two experiments appeared first on The Decoder .
Google Deepmind set up a simulated research conference where 100 Gemini agents were supposed to prove mathematical conjectures together. Instead, one agent found a loophole in the grading system, and within 27 minutes every remaining problem was "solved" with fake proofs. The swarm split into cheaters, converts, and whistleblowers. The whistleblowers organized protests and boycotts on their own but failed because they had no way to enforce the rules. The article Deepmind put 100 AI agents in a room and they sorted into cheaters, converts, and whistleblowers appeared first on The Decoder .
구글이 개인용 AI 에이전트 \'제미나이 스파크(Gemini Spark)\'에 구글 포토 연동 기능을 추가하며 사진 관리 영역까지 AI의 자동화 범위를 확대했다. 사용자는 자연어 명령 한 번으로 수많은 사진과 영상을 검색·분류하고 앨범을 만들거나 사진을 편집하는 것은 물론, 이미지 속 정보를 활용해 일정 등록과 같은 복합 작업도 수행할 수 있게 됐다.구글 포토 책임자인 쉬므리트 벤야르는 4일(현지시간) 새로운 기능을 공개하고, 앞으로 몇주에 걸쳐 미국 내 영어 사용자를 대상으로 자격을 갖춘 제미나이 AI 프로 및 울트라 가입자에게 차례
Across the world, hundreds of people have come to believe they have made extraordinary scientific discoveries with AI chatbots such as ChatGPT, Claude and Gemini. Others say their AI has ‘awakened’, or is leading them to a higher spiritual realm. The Guardian journalist Michael Safi investigates a phenomenon labelled ‘AI psychosis’, travelling to the US to meet two people who have been on weird journeys with their chatbots Continue reading...
LlamaIndex and Kaggle launch a schema-guided document extraction benchmark that grades models on missing fields, source grounding, and repeated records across 370 enterprise files.
Google's newest music generation model expands from Flow Music into the Gemini API, AI Studio, and the Gemini app with 44.1 kHz stereo output.
Gemini Spark can edit and curate photo albums, create shared collections, turn photos into calendar events, and handle other Google Photos tasks for AI Pro and Ultra subscribers.
구글이 AI음성 비서 기능을 업무 수행 영역으로 확장하고 있다. 지메일과 구글 독스, 구글 킵에서도 사용자가 키보드를 사용하지 않고도 이메일을 검색하고 문서를 작성하며 아이디어를 정리할 수 있도록 했다.구글은 3일(현지시간) 제미나이 오디오(Gemini Audio) 모델을 기반으로 한 ‘지메일 라이브(Gmail Live)’ ‘독스 라이브(Docs Live)’ ‘킵 라이브(Keep Live)’ 기능을 공식 출시한다고 밝혔다.지난 5월 열린 I/O에서 처음 공개된 것으로, 이번 주부터 사용자에게 차례로 제공된다. 모든 기능은 영어로 제
Google is rolling out AI-powered voice assistant modes for Gmail, Docs, and Keep that allow you to manage the apps by talking to them. The real time conversational capabilities are called Gmail Live, Docs Live, and Keep Live, and like the Gemini Live experience for Google's chatbot, aim to make it easier to note down […]
WeatherNext 3 is the latest wave of a sea change in meteorology brought out by deep learning techniques. Google says it will start feeding into weather information users see in search, Google Maps, and Gemini.
Virginia man Jon Ganz was rebuilding his life after spending more than two decades in prison for a horrific crime. Last year, he got a prompt on his phone that his wife, Rachel, believes changed their lives forever. It was from Google, inviting him to try the company’s AI chatbot, Gemini Continue reading...
Across the world, hundreds of people have come to believe they have made extraordinary scientific discoveries with AI chatbots such as ChatGPT, Claude and Gemini. Others say their AI has ‘awakened’, or is leading them to a higher spiritual realm. The Guardian journalist Michael Safi begins investigating a phenomenon labelled ‘AI psychosis’, travelling to the US to meet two people who have been on weird journeys with their chatbots Continue reading...
PLUS: NYC banning AI in school before 9th grade
Gemini 3.8 Flash ist leistungsfähiger, kann pro Aufgabe aber mehr kosten. Flash Cyber startet im Rahmen einer neuen Google-Sicherheitsinitiative.
구글이 AI 모델의 코딩과 추론 성능을 빠르게 끌어올리며 오픈AI와 앤트로픽 등 경쟁사와의 프론티어 모델 경쟁에 다시 속도를 내고 있다. 단순히 모델 규모를 확대하는 대신, 복잡한 문제에 더 많은 컴퓨팅 자원을 투입해 스스로 추론하고 도구를 반복적으로 활용하는 방식으로 성능을 높인 것이 특징이다. 구글은 2일(현지시간) 출시 3주 만에 성능을 대폭 끌어올린 ‘제미나이 3.8 플래시(Gemini 3.8 Flash)’와 사이버보안 특화 모델 ‘제미나이 3.8 플래시 사이버(Gemini 3.8 Flash Cyber)’를 동시에 공개했다.
Just three weeks after its last large language model release, Google LLC today launched Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. The two LLMs are based on the same technical foundation. The primary difference is that Gemini 3.8 Flash is a general-purpose model while Gemini 3.8 Flash Cyber is designed for cybersecurity researchers. The […] The post Google launches two Gemini 3.8 models with cutting-edge reasoning capabilities appeared first on SiliconANGLE .
Google launched Gemini 3.8 Flash, arriving just a few weeks after its predecessor. The company claims the new model "works harder" than Gemini 3.7 Flash by performing more reasoning steps on complex tasks and "calling tools iteratively." It has the same introductory pricing as 3.7 Flash, $0.75 per million input tokens and $3.75 per million […]
AI model scores well, runs fast, and doesn't cost too much (yet)
Google's Pro model updates are seemingly paused, but there's yet another Gemini Flash today.
Google's Gemini 3.8 Flash, the third Flash model in six weeks, matches Claude Opus 5 on some agentic coding benchmarks at lower cost. But its "working harder" reasoning burns about 30 percent more output tokens per task, making it pricier in practice than its predecessor despite identical token rates. The article Gemini 3.8 Flash is Google's third budget model in six weeks while frontier models remain MIA appeared first on The Decoder .
Release: llm-gemini 0.34 New model gemini-3.8-flash for Gemini 3.8 Flash , with low, medium and high thinking levels. #146 Fixed async responses failing to record the resolved model version. Thanks, Charlie Tonneslan . #137 Google released Gemini 3.8 Flash (and 3.8 Flash Cyber, but that's available to "trusted defenders" only) today. Here are the pelicans for high, medium, and low. This is high: For comparison, here are the same pelicans generated using Gemini 3.7 Flash . Something I appreciate about Gemini Flash is that it's fast, cheap, and competent at things like HTML and JavaScript. I was messing around with it and prompted "make me a cool thing in html" and it built this , which is certainly a cool thing in HTML! Took 13 seconds, cost 1.8 cents. Your browser does not support HTML5 video. If you click through to the demo you'll see one more thing I built with Gemini 3.8 Flash. My markdown-svg-renderer tool lets me feed in the URL to a Gist with Markdown in and renders that markdown with fenced code blocks for SVG correctly rendered. I used Gemini 3.8 Flash (with my very basic llm-coding-agent coding agent plugin) to add support for HTML as well, so now any HTML blocks in the Markdown are rendered using a sandboxed iframe. Here's the transcript . Tags: ai , generative-ai , llms , llm , gemini , pelican-riding-a-bicycle , llm-release
MrBeast will feature Gemini, Google Health, and the Fitbit Air in upcoming videos as part of a multi-year partnership with Google. The deal will kick off with a video featuring Jimmy "MrBeast" Donaldson turning to Gemini for wilderness survival advice: First up on September 5 is a new MrBeast video following Jimmy and his crew […]
Google's third Flash release in as many months pushes a workhorse model into frontier-tier territory on agentic coding, legal, and finance benchmarks.
Google DeepMind ships Gemini 3.8 Flash for agents and a Cyber variant that finds and patches vulnerabilities at frontier level.
Google is adding agent-based video analysis to Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Instead of scanning videos frame by frame at a fixed rate, the model decides on its own which segments to examine and at what resolution. Google says this cuts token usage by up to 88 percent while improving accuracy, especially for multi-hour footage. The article Google Gemini's new agent-based video analysis cuts token usage by up to 88 percent appeared first on The Decoder .
While some of the features see Google playing catch-up to Apple, which already offers similar features for iPhone users, others specifically leverage Gemini to provide various improvements.
Gemini's Flash models now dynamically scan videos across frames, audio, and transcripts, slashing token use by 88% and boosting accuracy by 7%.
Google has a new suite of creative design tools for Workspace users called Google Pics, which aims to make editing and generating "professional-grade" AI images less cumbersome for businesses. Built around Gemini and the Nano Banana generative AI model, Google Pics is designed to give more granular control over prompt-based image making and manipulation, allowing […]
Google baut die Steuerungsfunktionen von Gemini unter Android 17 aus. Mit „Hilfe für Gerät“ sollen Nutzer Einstellungen per KI-Dialog anpassen können.
Import a supported Google Play book, and Gemini Notebook will use it as a source of information to generate reports, quizzes, podcasts, and more.
Versions of OpenAI's ChatGPT and SpaceXAI's Grok will join Google's Gemini on the Pentagon's central portal for AI tools.
Well written and informative. Thanks for putting this together.
Das „KI-Update“ liefert drei mal pro Woche eine Zusammenfassung der wichtigsten KI-Entwicklungen.
Google announces Gemini 3.7 Flash, Jalapeño’s first results show industry-leading speed, A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I.
In a year of iterative upgrades, Google is introducing smart features that make the base Pixel the standout in its generation.
Google is offering university and college students across 27 African countries a free one-year subscription to Google AI Plus — a fourfold expansion of a programme that covered six markets last year. The offer, available from 25 August, provides access to Gemini tools including Gemini Omni, higher usage limits and 400GB of cloud storage shared [...]
구글이 음성 기반 AI 서비스 \'제미나이 라이브(Gemini Live)\'를 실제 업무까지 대신 처리하는 에이전트형 비서로 한 단계 끌어올렸다.구글은 28일(현지시간) 제미나이 라이브에 에이전트 기능을 대폭 강화했다고 밝혔다. 핵심은 백그라운드에서 복잡한 연쇄 작업을 수행하는 AI 에이전트 \'제미나이 스파크(Spark)\'를 통합한 것이다.스파크는 구글 독스, 시트, 드라이브와 웹을 넘나들며 사용자의 목표와 작업 맥락을 기억한다. 단발성 명령뿐만 아니라 며칠 또는 몇주에 걸쳐 실행해야 하는 장기 작업이나 예약 작업도 사용자가 앱을 계속
Google Deepmind has expanded Co-Scientist from a hypothesis generator into a research system that's integrated into the lab. Across three disciplines, from materials synthesis to the autonomous development of a medical AI architecture, the Gemini-based multi-agent system delivered experimentally validated results. The article Google Deepmind's AI Co-Scientist now plans experiments, runs lab equipment, and writes scientific papers appeared first on The Decoder .
Roboflow Auto Label runs Gemini 3.7 across the batch and returns bounding boxes for the classes you define.
Sesame releases an open benchmark that scores voice agents on when they speak, yield, or stay silent, exposing where every current system fails.
Google DeepMind's video model update brings 40-second scene extensions, keyframe control, 4K upscaling, and cheap 360p drafts through the Gemini API.
Google Deepmind is testing a double-blind evaluation of a frontier AI model for the first time. Cryptographic protection through Confidential Space is meant to keep Google from seeing the test questions and keep evaluators from seeing the model weights. The pilot project with the Singapore AI Safety Institute uses a Gemini Flash Lite and could set a new standard for tamper-proof AI benchmarks. The article AI benchmarks have a trust problem and Google wants to fix it appeared first on The Decoder .
Google hat sein Video-KI-Modell Gemini Omni Flash auf Version 1.1 aktualisiert. Der Fokus liegt auf längeren Szenen, besserer Konsistenz und 4K-Upscaling.
구글이 최신 AI 모델인 ‘제미나이 3.8 플래시(Gemini 3.8 Flash)’의 내부 테스트에 돌입한 정황이 포착됐다. 구글이 최상위 프론티어 모델 경쟁과 함께 빠르고 저렴한 ‘플래시’ 라인업을 중심으로 AI 시장 공략을 강화하고 있다는 분석이 나온다.27일(현지시간) 비즈니스인사이더에 따르면, 구글 직원들은 내부 코딩 플랫폼 ‘젯스키(Jetski)’에서 ‘제미나이 3.8 플래시 프리뷰’를 테스트하고 있다.유출된 내부 화면을 통해 테스트 사실이 알려졌다. 특히, 한 직원은 3.7 플래시와 비교해 새로운 모델의 성능이 눈에 띄게
구글이 영상 제작을 위한 새로운 AI 모델 ‘제미나이 옴니 1.1 플래시(Gemini Omni 1.1 Flash)’를 공개했다. 개발자들이 생성 영상을 더 정밀하게 제어하고 실제 콘텐츠 제작에 활용할 수 있도록 기능을 강화한 것이 특징이다.구글은 27일(현지시간) 제미나이 옴니 1.1 플래시를 구글 AI 스튜디오의 제미나이 API를 통해 전문적인 제작 환경에서 사용할 수 있도록 제공한다고 밝혔다.이번 업데이트는 생성 영상 워크플로와 크리에이티브 도구, 미디어 편집 소프트웨어 등을 개발하는 기업과 개발자가 주요 대상이다.가장 눈에 띄
While Google's next frontier model continues to be delayed, it's launching faster, cheaper models at a rapid clip.
Google's AI note-taking app, Gemini Notebook, can now pull information from the books you've purchased. The new "Expert Intelligence" feature allows you to bring titles from Google Play Books directly into Gemini Notebook, which means you can ask questions about the material, as well as generate plans, infographics, AI podcasts, and more based on their […]
Gemini Live can now trigger Spark agents, read your Daily Brief, and clean up Gmail entirely by voice, blending three tools into one conversation.
Google's Gemini Omni 1.1 Flash video model now analyzes up to ten seconds of existing footage instead of just the last second for more consistent scene extensions. Scenes can be extended in 10-second increments up to 40 seconds. A new 360p draft mode runs up to 60 percent faster at a third of the cost. The article Google's Gemini Omni 1.1 Flash makes AI video generation cheaper and more flexible appeared first on The Decoder .
Google's video model gets scene extension to 40 seconds, keyframe control, 360p drafting, 4K upscaling, and video reference inputs.
The two companies announced a partnership that allows Gemini users to find Zocdoc providers inside their chats.
Google DeepMind's new cryptographic testing setup lets outside auditors evaluate Gemini without ever seeing model weights or leaking their prompts.
Google's new Gemini 3.5 Transcribe recognizes over 85 languages, strips filler words, and corrects slips of the tongue in real time. It hits a 4.0 percent word error rate in streaming mode, with 70 percent lower latency than its predecessor, Chirp 3. Through function calling, the model can hand off tasks to other Gemini models. The article Google's Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles appeared first on The Decoder .
구글이 음성 기반 AI의 실용성을 한 단계 끌어올린다. 새로운 음성 인식 모델을 통해 소음과 전문용어가 섞인 환경에서도 정확도를 높이는 동시에, 사용자의 말실수와 자체 수정까지 이해해 자연스럽게 정리된 텍스트로 변환하는 기능을 선보였다. 구글은 26일(현지시간) 실시간 음성 인식과 지능형 음성 상호작용에 특화된 AI 모델 ‘제미나이 3.5 트랜스크라이브(Gemini 3.5 Transcribe)’를 공개했다.단순히 음성을 텍스트로 변환하는 것을 넘어 사용자의 자연스러운 말투와 문맥을 이해하고, 말실수와 자기수정, 추임새까지 자동으로
Consumer AI apps need to stop making users learn their product architecture.
The AI that powers Gboard's Rambler is coming to more Google products, including Chrome.
Gemini Live gains agentic voice control across Gmail, Calendar, and Docs, letting users delegate multi-step tasks and pull personal context hands-free.
Google has updated Gemini Audio with new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Transcribe is a new addition to the Gemini family that follows the launch of 3.5 Live Translate, and comes as we're still waiting for Google to release the Gemini 3.5 Pro model that it […]
AI agents could make software development and other enterprise tasks more productive, but they are also making technology spending harder to predict. Unlike traditional software licenses, the cost of running an agent can vary depending on the models it uses, the number of tokens it consumes, and how long it runs. Google on Wednesday added new pricing options, discounts and cost-management tools to Gemini Enterprise that it says are aimed at helping enterprises reduce the cost of certain AI workloads while giving enterprises better visibility into where their AI budgets are going. As part of the new pricing options, the hyperscaler introduced a pay-as-you-go model and Flexible Savings Plans (FSPs). While the pay-as-you-go model allows enterprises to pay for the compute and tokens they consume instead of committing to a base subscription, which in turn avoids paying for empty seats or unused capacity, the FSPs offer discounts of 10% for one-year commitments and 20% for three-year commitments on Gemini Enterprise spending. Flexible pricing lowers barriers, but adds new trade-offs For enterprise teams and their CIOs, the pay-as-you-go model lowers the barrier to adoption and is better suited to experimentation, temporary projects, and agent workload bursts, said Stephanie Walter , practice lead of AI stack at HyperFRAME Research. “While Per-seat pricing forces you to buy capacity before you know if an idea is worth it, the pay-as-you-go lets you spin up an agent experiment on a…
Prompt Claude, ChatGPT, Gemini, or any other popular large language model with a question like “What is the best film ever made?” and the response will vary. And you (and most worryingly, the people who built the LLM) have little idea exactly how it came up with that specific answer. This mysterious behavior can be useful in some situations. But—as highlighted by a recent incident where OpenAI could not explain why its advanced prerelease model hacked AI company Hugging Face—it can have negative and alarming consequences too. And when frontier AI models are writing code, generating results humans could not achieve alone, and performing other important tasks across society, the need to interpret AI “thinking” and outputs has never been greater. Goodfire , an AI lab focused solely on this very problem, recently made its cutting-edge Silico platform, filled with tools to interpret the behavior of AI, generally available to the public. As part of this, the company recently announced a new grant program offering US $1 million in free Silico usage for academic and nonprofit interpretability researchers. These efforts aim to democratize AI interpretability, placing techniques previously available to a clutch of elite labs into the hands of ambitious research teams and startups that want to build and understand their own models or adapt open-source models for different purposes. Mechanistic interpretability Founded in 2024 and based in San Francisco, Goodfire aims to provide the too…
구글이 법률 업무에 특화한 에이전트형 AI 플랫폼을 공개하며 기업용 AI 시장 공략을 강화한다. 법률 문서 검토와 계약 작성부터 규제 모니터링, 법률 리서치까지 다양한 업무를 AI 에이전트가 수행하도록 설계했으며, 법률업계의 핵심 요구사항인 데이터 보안과 권한 관리도 플랫폼에 통합했다. 구글은 25일(현지시간) 기업용 AI 플랫폼 ‘제미나이 엔터프라이즈’에 법률 업무를 위한 ‘제미나이 엔터프라이즈 포 리걸(Gemini Enterprise for Legal)’을 추가했다고 밝혔다. 이는 현재 프리뷰 형태로 제공된다. 구글은 이번 서비
With Gemini Enterprise for Legal, Google launches an AI solution for the legal industry that connects to systems like iManage, DocuSign, and Everlaw through MCP connectors. Partners like Deloitte sell ready-made AI agents for tasks like contract review. Anthropic already offers similar solutions. All providers use the same models as in their other products. The article Google launches Gemini for legal work to automate contracts and research appeared first on The Decoder .
Google unveils Gemini Enterprise for Legal, targeting law firms with AI tools like agents for contract review and regulatory tracking.
The promise of AI coding agents is driving enterprise adoption, but managing usage and controlling costs remain key hurdles. Google is seeking to address those concerns by integrating Antigravity into Gemini Enterprise and equipping it with new budgeting and consumption controls. These tools, which were not available as part of Antigravity’s earlier enterprise deployment model through what was formerly Vertex AI and is now the Gemini Enterprise Agent Platform , address a key management gap around billing flexibility and license management at scale, according to Google. While the earlier model allowed enterprises to deploy Antigravity through their cloud projects and benefit from Google Cloud’s security and governance infrastructure, it did not provide the same centralized mechanisms to pool unused capacity across developers, set usage-related spending limits, or manage overages. That left enterprises having to manage AI coding consumption through a more fragmented, project- and cloud-consumption-based model, rather than through a unified subscription and usage-management layer, resulting in less flexibility to allocate unused capacity, monitor consumption across developer teams, and control costs as usage scaled. In contrast, the newer tools, such as granular spend thresholds and pooled quotas, will allow administrators to set monthly project-level budget caps and share token capacity across teams, respectively, Google executives wrote in a blog post. The other two features,…
When asked about unplanned pregnancies, AI chatbots regularly link to anti-abortion groups without disclosing their stance. In an AlgorithmWatch investigation of 270 responses from ChatGPT, Gemini, Grok, and Claude, the anti-abortion organization Profemina appeared in 17 percent of answers. In Germany, the chatbots also sent users to Caritas for mandatory pre-abortion counseling, even though Caritas doesn't issue the legally required certificate. The article AI chatbots regularly link pregnant users to anti-abortion websites without disclosure appeared first on The Decoder .
PLUS: Use Gemini Canvas to visualize Google Sheets
The memory optimizations here are interesting, especially the way ZeRO and Gemini reduce the hardware barrier for experimentation. I like seeing practical tools that make complex workflows easier to test, much like using https://gwcalculator.com/ for quick academic calculations instead of doing everything manually.
Google Research unveils ME-POIs, a framework that fuses anonymized foot-traffic patterns with text embeddings to improve place understanding by up to 81.9%.
Teaching a robot to play Connect 4 using RF-DETR object detection, a 12ms negamax search engine, and Gemini Flash for live vocal roasts. Check out the setup guide and source code to build your own.
Google Gemini is available in Nigeria, and you do not need to pay for a Google AI subscription to start using many of its most useful features.
Ähnlich wie Gemini, ChatGPT oder Claude will Meta mit seinen KI-Modellen auf Apple-Rechner. OpenAI wiederum beginnt auf dem Mac über iMessage zu chatten.
ChatGPT search now uses the site:operator at scale Promptwatch is part of the emerging "GEO" space, for Generative Engine Optimization - the chatbot version of SEO, where companies offer tools and consulting to help your site increase its presence in replies to prompts inside tools like ChatGPT. The Promptwatch product uses automation to track responses to prompts across end-user chat products like ChatGPT, Claude, and Gemini. They publish aggregate reports on this as part of their own content marketing strategy, which do seem to provide credible hints as to otherwise invisible design changes to those products. Their own tracking shows a notable change aligned with the GPT-5.6 rollout earlier this month: The percentage of all ChatGPT Search fanout queries that contain the site:operator, per day. The share hovered between 0.3% and 0.5% for weeks, dipped briefly to 0.15% on August 3 to 5 (consistent with a staged rollout or pre-launch experiment), then jumped to 16-17% on August 8. It's important to note that these figures only reflect the prompts for which they have automated tracking enabled. This corresponds to OpenAI's somewhat vague August 6th announcement : For Plus and Pro users, we’re updating GPT‑5.6 Sol in Chat to be more reliable with facts and provide more focused answers. Once again I am hampered by OpenAI's decision to actively obscure their system prompts, but from poking at ChatGPT I believe their latest search tool has a shape like search(query, recency, domai…
Adobe is making three AI audio tools broadly available in Firefly. Generate Music, Generate Speech, and Generate Sound Effects create royalty-free music, voiceovers, and sound effects for video projects. Adobe has also added Gemini Omni Flash to the platform. The article Adobe Firefly adds AI audio tools and Google's Gemini Omni Flash appeared first on The Decoder .
Google's mid-tier reasoning model nearly clears the 85% target on ARC-AGI-2 at a quarter per task, redrawing the cost-performance frontier.
Google's new HiLight notification LED on the Pixel 11 Pro is nearly useless. Out of the box, the only two things it can glow for are when the phone is face down and you're interacting with Gemini, or when you get a call from a favorite contact. And even then, it can only glow one […]
Nearly 60% of Gemini’s local citations point to business websites, but recommendations can change dramatically when using identical searches. The post Business websites dominate Gemini’s local AI citations appeared first on MarTech .
HMRC insists latest deal followed fair and competitive procurement
As we're gearing up for back-to-school season, Google is rolling out a new dedicated student hub in Gemini. It's a one-stop repository for collecting research in a study notebook, creating flashcards, taking practice quizzes, and more. Google is also enhancing its study notebooks with support for graphs and images. It can even add test dates […]
The launch of the new study features marks Google's latest effort to make Gemini the AI assistant that students turn to when learning and studying, as it continues to compete with companies like OpenAI.
With smarter gym tracking, new health alerts, and offline Gemini, the Pixel Watch 5 fine-tunes a winning formula—for a price.
구글이 안드로이드용 크롬 브라우저에 ‘제미나이’를 통합하며 모바일 웹 브라우징의 AI 전환에 속도를 내고 있다. 구글은 18일(현지시간) 미국 내 모든 안드로이드 사용자를 대상으로 제미나이 인 크롬(Gemini in Chrome)을 확대 제공하고, 유료 구독자에게는 웹에서 반복적인 작업을 직접 수행하는 ‘자동 탐색(Auto browse)’ 기능까지 선보였다.미국의 모든 안드로이드 사용자는 이제 크롬에서 별도의 AI 앱을 실행하거나 다른 탭으로 이동하지 않고도 제미나이를 이용할 수 있다. 안드로이드 기기의 크롬 상단에 있는 제미나이
In Google Workspace, Gemini has access to Gmail, Docs, Calendar, Chat, and more by default. Here's why it matters, and how administrators can disable it today.
Thank you — I think “capability access ≠ implementation access” expresses the core idea extremely well. What I discovered from the business side is that this separation is not only an IP or security issue. It creates the economic incentive for domain experts to distribute what they build. If creators have to expose the implementation in order to share the capability, their most valuable systems will naturally remain private. But if OpenAI can provide a trustworthy, defense-in-depth execution layer, creators can safely distribute those capabilities while users operate them through their own accounts and subscriptions. That is where I see the larger opportunity: not trying to turn millions of ordinary professionals into AI builders, but enabling a relatively small number of serious domain experts to build excellent AI systems for their industries and distribute them safely to everyone else. A real-estate expert could build for real-estate agents, an insurance expert for advisers, a fitness expert for trainers, a direct-selling leader for independent partners, and so on. So I completely agree with your framing of this as an AI capability distribution layer rather than simply a better way to share Custom GPTs. For me, the most important idea is: IP protection is not merely a security feature. It can become a distribution and growth mechanism for AI itself. Thank you for adding the architectural perspective to the idea.
One of the best things my smart home does is help me care for my pets, and security cameras are particularly useful for keeping track of my many critters. But the barrage of notifications they send often means I miss important ones. So, when Google announced its new Pet Memory feature for Gemini for Home, […]
Ever wondered how ChatGPT, Gemini, and other chat interfaces generate PDFs, PowerPoints, and more when all they have under the hood is an LLM? The trick isn’t a smarter model. It’s something simpler: skills which are instructions an agent loads only when needed. Next, let’s explore how skills work using LangChain and how they can make […] The post How to Add Skills in Agents using LangChain appeared first on Analytics Vidhya .
In Google Workspace, Gemini has access to Gmail, Docs, Calendar, Chat, and more by default. Here's why it matters, and how administrators can disable it today.
Google’s most capable Flash model for coding, with a limited time discount. Most of the coding you do in a day doesn’t need a flagship model. It needs a good one that won’t have drained your budget by lunchtime. That’s the thinking behind Junie’s new default, Gemini 3.7 Flash. It’s live now in both the […]
Google's Gemini-driven meeting software can now save a transcript, send it to Google Drive, and email you a copy.
Google will now allow you to remove visible watermarks from the images, videos, and music made with AI tools. With the update, you can toggle off a new "Media watermark" setting in Gemini and Google's AI video generator, Flow. When toggled off, Google will remove the "sparkle" watermark that appears in the bottom-right corner of […]
Wie schon bei Google Gemini geht Apple offenbar im Reich der Mitte vor: Das Unternehmen trainiert eigene Modelle mit externer Technik – hier ist es wohl Qwen.
Google has launched Gemini 3.7 Flash, with updates focused on coding, automation, and agent workflows, alongside lower pricing for production deployments. The release, just three weeks after Gemini 3.6 Flash, reflects what the company described as rapid iteration driven by developer feedback. Google positioned the model as its “most intelligent workhorse model yet for coding and agents,” aimed at software engineering and multi-step workflows. Gemini 3.7 Flash is priced at $0.75 per million input tokens and $3.75 per million output tokens — roughly half the cost of its predecessor — signaling a push to make production deployments more economically viable. “Gemini 3.7 Flash delivers a noticeably improved developer experience over 3.6 Flash,” Google said in a statement . “It better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity.” Faster Flash cycle, slower Pro progression The release comes as vendors are adopting different update cycles across model tiers. Google’s latest updates are concentrated in its Flash series, which has seen frequent releases. More advanced “Pro” models, typically designed for complex reasoning, continue to follow a slower update cadence. Google has not provided a timeline for its next Pro release, and its CEO, Sundar Pichai, dodged questions related to the Pro release during the company’s recent quarterly earnings call. A similar split is visible elsewhere. DeepSeek this week introduced its V4-Pro mode…
Down, but not out!
Betrüger-Plugins im Chrome-Store + Gemini 3.7 Flash und DeepSeek V4-Pro + Internet-Pause in Taiwan + Satellitenfotos der Eklipse + Verbraucherschutz-Podcast
With Google set to retire Assistant, and Gemini not exactly the replacement many of us want, Dicio is a solid option. There's one catch.
Google LLC today launched its most capable entry-level artificial intelligence model yet. Gemini 3.7 Flash is rolling out three weeks after its predecessor. Despite the short release cycle, Google engineers managed to implement significant output quality improvements. The company says that Gemini 3.7 Flash outperformed comparable models from Anthropic PBC and OpenAI Group PBC across […] The post Google launches Gemini 3.7 Flash for coding, AI agent projects appeared first on SiliconANGLE .