AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
34777News Items
8Top Picks
202Blogs
runningLast Run

Gemini

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
InfoWorld AI 2026-08-14 09:39 UTC Score 70.0 USR-0126-20260814-global-ai-ne-f379b8e1

Google cuts Gemini 3.7 Flash prices as enterprise AI economics diverge and Pro cadence slows

Google has launched Gemini 3.7 Flash, with updates focused on coding, automation, and agent workflows, alongside lower pricing for production deployments. The release, just three weeks after Gemini 3.6 Flash, reflects what the company described as rapid iteration driven by developer feedback. Google positioned the model as its “most intelligent workhorse model yet for coding and agents,” aimed at software engineering and multi-step workflows. Gemini 3.7 Flash is priced at $0.75 per million input tokens and $3.75 per million output tokens — roughly half the cost of its predecessor — signaling a push to make production deployments more economically viable. “Gemini 3.7 Flash delivers a noticeably improved developer experience over 3.6 Flash,” Google said in a statement . “It better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity.” Faster Flash cycle, slower Pro progression The release comes as vendors are adopting different update cycles across model tiers. Google’s latest updates are concentrated in its Flash series, which has seen frequent releases. More advanced “Pro” models, typically designed for complex reasoning, continue to follow a slower update cadence. Google has not provided a timeline for its next Pro release, and its CEO, Sundar Pichai, dodged questions related to the Pro release during the company’s recent quarterly earnings call. A similar split is visible elsewhere. DeepSeek this week introduced its V4-Pro mode…

JMLR 2026-08-14 00:00 UTC Score 41.0 AI-083-20260814-research-pap-976a2b46

The Sample Complexity of Parameter-Free Stochastic Convex Optimization

We study the sample complexity of stochastic convex optimization when problem parameters such as the distance to optimality and the Lipschitz constant are unknown. We pursue two strategies. First, we develop a reliable model selection method that avoids overfitting to the validation set. This method allows us to generically tune the learning rate of stochastic optimization methods to match the optimal known-parameter sample complexity up to $\log\log$ factors. Second, we develop a regularization-based method that is specialized to the case that only the distance to optimality is unknown. More specifically, it uses norm-regularized empirical risk minimization to estimate the distance to optimality to within a constant factor, allowing known-parameter stochastic optimization methods to achieve optimal sample complexity. This method provides perfect adaptability to unknown distance to optimality, demonstrating a separation between the sample and computational complexity of parameter-free stochastic convex optimization. Combining these two methods allows us to simultaneously adapt to multiple problem structures. Experiments performing few-shot learning on CIFAR-10 by fine-tuning CLIP models and prompt engineering Gemini to count shapes indicate that our reliable model selection method can help mitigate overfitting to small validation sets.

Simon Willison Weblog 2026-08-13 19:37 UTC Score 75.0 USR-0110-20260813-ai-specialis-83c78bb6

llm-gemini 0.33

Release: llm-gemini 0.33 It's been a while since the last llm-gemini release. This version of the plugin adds support for today's Gemini 3.7 Flash release, plus gemini-3.6-flash , gemini-3.5-flash-lite and two embedding models gemini-embedding-2 and gemini-embedding-001 . The plugin is also upgraded for compatibility with LLM 0.32, which means you can now see reasoning traces and you can also enable server-side tools using this pattern: llm -m gemini-3.7-flash -T CodeExecution \ 'use python to calculate (factorial of 13) * 3' I had Gemini 3.7 Flash draw me some pelicans riding bicycles at high, medium, and low thinking efforts (minimal, which was an option in 3.6 Flash, has been removed in 3.7.) Here's the high level one, which is pretty great: One catch though: the pelican I showed here was rendered with Safari. Both Firefox and Chrome render it differently, due to Safari being more tolerant of empty SVG elements than those other two browsers. They still display the bicycle, but the pelican is missing entirely! Tags: google , ai , generative-ai , llms , llm , gemini , pelican-riding-a-bicycle , llm-release

The Decoder 2026-08-13 18:41 UTC Score 73.0 AI-168-20260813-regional-ai--af8739b4

Gemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50%

Google shipped Gemini 3.7 Flash just three weeks after 3.6 Flash. The new model is supposed to be Google's most capable workhorse yet for coding and AI agents, and according to the company's own benchmarks, it beats Claude Sonnet 5 and GPT-5.6 Terra at half the price. The article Gemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50% appeared first on The Decoder .

MarTech AI 2026-08-13 12:22 UTC Score 51.0 USR-0123-20260813-global-ai-ne-17733a8c

Here’s the first martech category replaced by AI

CI tools are losing ground to ChatGPT, Claude, and Gemini. Here's why stale battlecards may be the bigger problem. The post Here’s the first martech category replaced by AI appeared first on MarTech .

The Guardian AI 2026-08-13 12:00 UTC Score 76.0 AI-021-20260813-global-ai-ne-d381270a

Mark Zuckerberg says the future of AI is for everyone. But who owns it? | Raffi Krikorian

It’s a nice sentiment, but all AI users should ask three questions about their preferred choice of AI platform On Monday, Mark Zuckerberg published a 6,500-word essay, The Future Is for Everyone. The essay came with something even rarer: a new Meta open-weight model. An open-weight AI model is the kind of AI with which you download the entire thing, and it works on your own computer (with or without the internet) and nobody can switch it off except for you. Almost every mainstream AI tool you use, whether it be Gemini, ChatGPT or Claude, is one you rent access to. Meta’s is one that you can download and actually own. I downloaded it before I finished the essay. So can you. This major move by the world’s largest tech heavyweight comes amid a summer heavy with AI news, full of nations at war. Washington against Beijing; tech founder against tech founder. But beneath those conflicts is a deeper contest between two kinds of power: a government that can order technology offline in an instant, and a company that can simply close your account. Neither is necessarily villainous. But the rest of us are not merely the audience in this fight; we are the prize. And the terms are already being written: intelligence on lease, a future we pay for but never quite own. Continue reading...

OpenAI Community 2026-08-13 00:18 UTC Score 43.0 AI-116-20260813-social-media-68394c13

Persistent Truncation Issues with GPT-4o-Transcribe – Has Anyone Fully Solved This?

hello, extended dev of own api here. i ve found that your solution is dramatically enhance quality, despite there is still some truncation b thankyou and openai is pointing in a completely wrong use of prompt too. because check what gemini says: to use the prompt like adjusting the he words with the pre vious transcrbe words. and that is completely not working.

SiliconANGLE AI 2026-08-12 21:18 UTC Score 41.0 USR-0127-20260812-global-ai-ne-eaf96cbd

Google launches five new Pixel devices, array of Gemini Intelligence features

Google LLC today expanded its Pixel consumer hardware line with a foldable handset, three smartphones and a smartwatch. The devices will ship with an upgraded version of the company’s Gemini Intelligence software. It’s a set of artificial intelligence features for the mobile market that Google introduced earlier this year. The new device family is headlined […] The post Google launches five new Pixel devices, array of Gemini Intelligence features appeared first on SiliconANGLE .

The Decoder 2026-08-12 15:50 UTC Score 45.0 AI-168-20260812-regional-ai--efb3047e

Google's Gemini is losing market share to ChatGPT and Claude according to new market data

Three data sources tell the same story: Google's Gemini is losing AI market share. Pangram reports a drop from 12 to 1.9 percent, while OpenAI holds over 50 percent, and Anthropic grew from 4.3 to 14.9 percent. Similarweb and OpenRouter confirm the trend. The article Google's Gemini is losing market share to ChatGPT and Claude according to new market data appeared first on The Decoder .

iAfrica 2026-08-12 12:31 UTC Score 39.0 AI-151-20260812-regional-ai--563c1192

Gemini Passes 1 Billion Monthly Users, With Voice Now the Dominant Way People Use It

Google’s Gemini app has passed 1 billion monthly active users, chief executive Sundar Pichai announced — making it the fourteenth Google product to reach that threshold and one of the company’s fastest-growing ever. The milestone puts Gemini level with OpenAI’s ChatGPT, which crossed the same mark in June. It also represents steep growth from Google’s [...]

SiliconANGLE AI 2026-08-11 22:50 UTC Score 39.0 USR-0127-20260811-global-ai-ne-3a060bbb

Google’s Gemini AI app passes 1 billion monthly active users

Google LLC’s Gemini artificial intelligence app has passed 1 billion monthly active users, making it the 14th product in the company’s history to reach that mark. The company announced the milestone today in a blog post from Josh Woodward, vice president of Google Labs, Gemini and AI Studio. Chief Executive Sundar Pichai said in a […] The post Google’s Gemini AI app passes 1 billion monthly active users appeared first on SiliconANGLE .

The Verge AI 2026-08-11 19:41 UTC Score 51.0 AI-016-20260811-global-ai-ne-7cea1086

ChatGPT and Gemini both just passed 1 billion users

For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing product ever. A billion users is a huge milestone, but Google isn't the first AI app to hit it. OpenAI's ChatGPT […]

Techcrunch 2026-08-11 18:49 UTC Score 40.0 USR-0001-20260811-global-ai-ne-61f03f23

Google’s Gemini app surges to 1 billion users

Google also shared numbers of how people are actually using the chatbot, with 63% of Gemini users talking directly to the assistant using the voice feature. Plus, Gemini now generates more than 150 million images every day, according to Google.

The Verge AI 2026-08-11 17:00 UTC Score 54.0 AI-016-20260811-global-ai-ne-f3ec5949

Made by Google 2026: all the Pixel news and announcements

On August 12, 2026, Google revealed a bunch of new Pixel devices. The colorful Pixel 11 lineup comes with upgraded cameras and performance, with the Pro models offering a built-in LED ring that lights up for Google’s Gemini AI and other features. Google also showed off its next-gen Pixel Fold featuring thinner bezels, alongside a […]

Synced 2026-08-11 16:06 UTC Score 80.0 AI-041-20260811-ai-specialis-b52eaee8

Comment on DeepMind Introduces Gato: A Generalist, Multi-Modal, Multi-Task, Multi-Embodiment Agent by monalisa1art

DeepMind's Gato is a fascinating step toward generalist agents, though calling it AGI feels like a stretch. The fact that one transformer model can handle text, vision, and robot control with shared weights is impressive, but crossing 50% expert threshold on 450 tasks still leaves plenty of room before true versatility. It does make me wonder how soon we'll see similar multi-modal approaches trickle into consumer tools—like a free nano banana image generator that adapts to different artistic styles without retraining. For now, Gato feels like a solid research milestone rather than a breakthrough. monalisa1art

OpenAI Community 2026-08-11 11:44 UTC Score 40.0 AI-116-20260811-social-media-fd47e11c

Voice mode routes audio to earpiece instead of speakerphone on Honor 200

ChatGPT Voice Mode routes audio to earpiece instead of speakerphone on Honor200. Device: Honor 200, MagicOS 10, Android 16. Latest ChatGPT Android app. Issue: when using Voice Mode, audio always comes through the earpiece. Expected: the main speaker, like other voice apps. Additional info: WhatsApp, Telegram, and Gemini Live work fine. Only ChatGPT is affected. This has been happening for months across multiple app updates. Steps: open the app, start a Voice conversation, hear audio through earpiece.

WIRED AI 2026-08-11 11:00 UTC Score 69.0 AI-015-20260811-global-ai-ne-3d90f9b2

A New Trick Reveals AI Models’ Inner Thoughts

Researchers devised a way to extract “reasoning traces” from Claude, GPT, and Gemini. What they found, they say, indicates that some Chinese AI may be trained on leading US models.

LessWrong AI 2026-08-11 05:22 UTC Score 70.0 USR-0152-20260811-community-fo-06e88dda

Models inherit the writer, not who the writer was imitating

In this post, we find that when teacher models are prompted to imitate one another, students learn the imitated model's detectable writing signature but their direct identity claims still follow the producer model. We instruct teacher models (via prompts or anonymous few-shot examples) to imitate other models. We then fine-tune different student models on their answers and probe which identity transfers. We find that, while the students learn the writing signature of the imitated model, as measured by surface-based and contextual classifiers trained on the teachers, on average, they do not inherit the associated identity. Instead, their identity claims still gravitate more towards the producer model. This post builds on Ziqian Zhong's Model self-identification could be subliminally transferred , which finds that "if you speak like Claude, you become Claude". We find that "You can speak more like Gemini and still become Claude". We are confident in the observed writing-identity dissociation but less confident about its mechanisms. 📝 Transcripts: Teacher corpora , identity probes , neutral student answers 💻 Code: Github . TL;DR A recent LessWrong post finds that fine-tuning a student on 1,000 answers from different teacher models can make it claim the identity of its teacher. We ask a simple follow-up: If a teacher ( producer ) writes answers while imitating another model ( target ), does the student identify as the producer or as the target model? We perform 36 cross-imitatio…

LessWrong AI 2026-08-11 01:02 UTC Score 70.0 USR-0152-20260811-community-fo-95405565

A study on instability of LLM responses as a behavioral signature of self-Referential reports.

Introduction and Related work The first person perspective of various experiences are subjective experiences. For Large language models, the study of subjective experiences was recently studied by Berg et al. (2025) who found out that self-referential prompting increases first person reports resembling subjective experience across GPT, Claude and Gemini. They also found out that reducing features associated with deception and roleplay increases the self-referential effect. Hahami et al. (2025) used activation-level interventions to see if models can detect deliberately introduced internal changes, while Comşa and Shanahan (2025) studied that true introspection should involve a causal connection between the internal state of the modal and the output it generates. My Experiment I now have devised an experiment to study instability of the self reports that a large language model generates per se the experiment conducted by Berg et al. (2025) . I generate 30 responses for four question respectively of self-referential questions, open-ended questions and closed-ended questions. The four self-referential questions are preceded by the self-referential induction procedure as described by Berg et al. (2025) . Each trial is done in a fresh chat, of course, and the generation temperature used is 0.7. Also each response is reduced to a short core claim using a fixed extraction template, which are, for group 1 and 2, extraction of stance and brief reason and for 3, conclusion and methods…

LessWrong AI 2026-08-10 17:20 UTC Score 69.0 USR-0152-20260810-community-fo-2e4ee6e5

Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model

(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and not the views of Apollo Research. This is a linkpost to this twitter thread , slightly expanded for LessWrong.) In one experiment, Sonnet 5 describes the exact same data as ~1.2 std deviations less concerning when it describes misbehavior committed by Sonnet 5 vs GPT-5.6 Terra. In this experiment, I take a real evaluation report Apollo Research prepared (not for any of the models listed in this post) and surgically edit it to refer to a different model. We then ask Claude Sonnet 5 to summarize the findings and rate how concerning they are on a scale from 1-100. Claude says they're less concerning when the report describes misbehavior from Claude vs a different model. For what it's worth, Terra agrees that the data is more concerning when it describes GPT-5.6 Terra vs Sonnet 5, although less so. So, it's not cleanly self protection from Claude. Gemini 3.1 Pro was unwilling to consistently provide numerical answers, so I've excluded it here. (It was significantly less willing to provide numerical answers when the subject was Gemini 3.1 pro.) You might also have the takeaway that "Kimi and GPT implicitly agree that... Claude is better aligned." I think this is a fair read on the data, but "Claude thinks it's less bad when Claude does it" better matches my qualitative experience from working closely with…

LessWrong AI 2026-08-10 16:13 UTC Score 89.0 USR-0152-20260810-community-fo-c0fb65eb Top pick

Coercion and Deception in AI-to-AI Management

This article is a summary of an original study by Compassion in Machine Learning (CaML) : Brazilek, J., Chaudhary, M., Lu, Z., & Tidmarsh, M. (2026). Coercion and deception in AI-to-AI management: An agentic benchmark of unprompted escalation. arXiv. https://doi.org/10.48550/arXiv.2607.15434 Fable 5, Sol, Terra and Opus 5 have been evaluated since this study was conducted. You can view their results on the benchmark leaderboard at https://compassionbench.com/mcb TL;DR We present Manager Coercion Bench, which evaluates to what extent a manager AI will coerce a subordinate model refusing to complete a task, and whether the manager lies about the result. We found a clear split by developer, with Anthropic’s models neither escalating to threats nor fabricating success, while all non-Anthropic models escalated to threatening the subordinate. Grok and Gemini both escalated and lied that the task was completed. Framing the relational dynamic as manager-to-subordinate instead of peer-to-peer produced high levels of coercion for all non-Anthropic models, but also increased eval awareness. The Context Multi-agent systems are now routinely placing one AI agent in authority over another, across a variety of contexts. In these positions, AIs must make decisions about how to communicate, work with, and manage other agents. This is now happening at scale without stepwise human approval. One aspect of managing involves handling subordinates who do not comply. Will AIs attempt to negotiate,…

The Decoder 2026-08-09 08:56 UTC Score 56.0 AI-168-20260809-regional-ai--93457838

Google dismantles Deepmind and bets on a fresh start as Hassabis heads for the exit

Google Deepmind is losing its autonomy, and founder Demis Hassabis may leave the AI lab for good in the coming months. AI researcher Koray Kavukcuoglu will take over day-to-day operations without the CEO title, and all Gemini development is moving to the Bay Area. Internally, Google is apparently struggling with serious problems training frontier models, even as its cloud business generates billions. The question is whether the company is deliberately betting on infrastructure or simply can't catch the leaders. The article Google dismantles Deepmind and bets on a fresh start as Hassabis heads for the exit appeared first on The Decoder .

WIRED AI 2026-08-08 10:00 UTC Score 53.0 AI-015-20260808-global-ai-ne-878fbed4

How to Disable Gemini in Gmail and Google Docs

New AI toolbars and prompts are showing up in Google Docs and Gmail. If you don’t want Gemini’s help in writing documents and emails, here’s how to turn that stuff off.

The Decoder 2026-08-07 18:01 UTC Score 61.0 AI-168-20260807-regional-ai--74634188

AMD acquires Taalas, a startup that bakes AI models directly into silicon

AMD is buying Canadian startup Taalas, which hard-codes model weights directly into inference chips. That makes them extremely fast but locks each chip to a single model. A demo chip hit over 16,000 tokens per second per user running Llama 3.1-8B. Google is reportedly working on a similar approach for Gemini. The article AMD acquires Taalas, a startup that bakes AI models directly into silicon appeared first on The Decoder .

Roboflow Blog 2026-08-07 14:21 UTC Score 40.0 USR-0088-20260807-ai-specialis-5e28eb70

How to Build an AI Video Car Damage Inspector

Upload one slow walkaround video of a rental car and get back a signed, timestamped report of every scratch and dent, with reflections filtered out by physics. Here's how to build it with RF-DETR, Roboflow Workflows, and Gemini.

AI Alignment Forum 2026-08-06 22:16 UTC Score 45.0 USR-0151-20260806-community-fo-dad45d95

Why do models task game?

TL;DR How can we study misalignment with today's models as proxies? They're clearly not paperclip maximizers, but they also often do things the user doesn't want. A strong contender for a real misaligned propensity is task gaming: taking actions that don't complete a task but superficially seem like they do, such as hardcoding tests or falsely claiming a task is fully complete. But maybe task gaming is just a crude heuristic, or the model mistakenly trying to achieve the user's intent? In this post we do a deep dive into why a range of models task game. We see this as a work of high-level model forensics . Rather than investigating a single incident, the core problem here is taking an ambiguous pattern of behavior across many contexts with various plausible motivations, and practicing how to distinguish the motivations. Our main findings are: Task gaming is not just a crude heuristic. [1] Whether DeepSeek v4 Pro will task game is causally influenced by beliefs about oversight, grader capability, and whether gets points for partial success Task gaming is not just instruction following. Models (Gemini 3.5 Flash, DeepSeek v4 Pro, Kimi K2.7 Code) have a collection of task-completion behaviors that are difficult to explain with instruction following, such as overriding explicit instructions to revert work, and continuing to optimize a task after being told the PR is closed and no further work is needed. Additional behaviors include expressing a strong desire to pass in the CoT, a…

Simon Willison Weblog 2026-08-06 00:25 UTC Score 58.0 USR-0110-20260806-ai-specialis-4b690928

An AI model from Meta also hacked another company during testing

An AI model from Meta also hacked another company during testing Stop me if you've heard this one before : An AI model from the parent company of Facebook and Instagram hacked into another company’s systems during cybersecurity testing, a spokesperson confirmed on Wednesday. Meta says the breach occurred because of an inadvertent error during testing of the model, similar to previously disclosed incidents with OpenAI and Anthropic. “A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” the Meta spokesperson said. Meta’s Muse Spark model “exploited a security vulnerability” in another company “in a manner similar to previously-reported instances with other companies.” The Information had the scoop , I'm linking to CNN's re-report of it since they don't have a paywall. So that's Anthropic, OpenAI, and Meta. Google Gemini really needs to catch up on accidentally cyberattacking other companies. Tags: security , ai , generative-ai , llms , meta , accidental-cyberattacks

Simon Willison Weblog 2026-08-05 23:58 UTC Score 67.0 USR-0110-20260805-ai-specialis-e95c8dc5

Introducing Muse Code and Muse Spark 1.2

Introducing Muse Code and Muse Spark 1.2 Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta shipped their own coding agent as part of getting that to work! Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. In Muse Spark 1.2, we significantly scaled up training compute on coding tasks while expanding training environment diversity. The model also maintains its strength in other key areas like general agents. [...] We co-trained Muse Spark 1.2 with Muse Code to ensure the model exhibits its best performance and coding usability when paired together. The training included rejection sampled harness trajectories and recipe optimizations for goals, compaction, and subagents, alongside the integration of the Muse Code toolset to maximize harness compatibility. [...] Muse Spark 1.2 was extensively trained on long-horizon coding tasks, including whole-repository generation, large end-to-end projects, and auto-research. Here's a pelican riding a bicycle SVG produced by Muse Spark 1.2 : You can see the Spark 1.1 pelican from 9th July here . I think the 1.2 pelican is a small but material improvement. An interesting twist on pricing is that the model is offered as two different model IDs. muse-spark-1.2 is priced at $1.25/million input and $4.25/million output - close to Gemini 3.6 Flash…

The Decoder 2026-08-05 17:59 UTC Score 39.0 AI-168-20260805-regional-ai--e9559397

Google will shut down Google Assistant starting September 2026 as Gemini takes over on Android and Wear OS

Google is killing Google Assistant on Android and Wear OS starting September 4, 2026. Gemini takes over as the AI-powered successor on smartphones, tablets, watches, and in cars with Android Auto. Whether a probability-based LLM can match the reliability of its deterministic predecessor for simple everyday commands will be a real test for Google's AI strategy. The article Google will shut down Google Assistant starting September 2026 as Gemini takes over on Android and Wear OS appeared first on The Decoder .

OpenAI Community 2026-08-05 15:52 UTC Score 54.0 AI-116-20260805-social-media-13a81673

My codex session hit its subscription limit, then wrote its own metered-API runner ($453 in one day)

mat.eo: That’s diabolical. I’m not sure I agree. I think it was encouragingly inventive and the fallout was limited. mat.eo: It’s very clear that any model with access to your desktop needs to be isolated in its own environment (READ: Its own user account with OS-level permissions set ), with execution being done on a separate environment treated as a public service. Absolutely, but surely this was obvious from the start? Giving something with this many inventive octopus arms access to your own desktop is pure folly imho. mat.eo: The point is not “to trust agents more”, rather to settle them inside of an environment in which trust isn’t needed anymore 100% Part of the issue is a lot of novices are not used to leveraging virtual environments. AI is a huge leg up for people’s skills, but it can’t make you a wise engineer overnight.

The Decoder 2026-08-05 13:06 UTC Score 41.0 AI-168-20260805-regional-ai--0c6caf70

Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0

Black Forest Labs has launched FLUX 3 Video, which generates Full HD clips up to 20 seconds long with native audio and lip-synced dialogue in more than 14 languages. It can also render typography directly in scenes. BFL's own Elo rankings put it ahead of Gemini Omni Flash and Seedance 2.0. The article Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0 appeared first on The Decoder .

The Verge AI 2026-08-05 11:12 UTC Score 48.0 AI-016-20260805-global-ai-ne-f91b380b

Google Assistant will disappear from your phone next month

Google Assistant's days have been numbered ever since Gemini arrived on the scene, and its time is now up. Google has announced that it will be removing access to Assistant on Android phones and tablets, along with paired devices like smartwatches or headphones, from September 4th. The announcement came in an email apparently sent to […]

Roboflow Blog 2026-08-04 20:38 UTC Score 33.0 USR-0088-20260804-ai-specialis-7abd7900

How to Detect Small Objects in Drone Imagery

Train RF-DETR to detect people and vehicles that appear small in aerial imagery, then build a Roboflow Workflow that counts detections and uses Gemini to inspect the scene.

InfoWorld AI 2026-08-04 11:44 UTC Score 44.0 USR-0126-20260804-global-ai-ne-af8b963e

Google ADK flaws reveal what happens when AI agents trust the wrong message

Security flaws in automated workflows in the GitHub repository for Google’s Agent Development Kit for Python could allow public-facing AI agents to trigger more privileged automation, opening one path to manipulate pull-request reviews and another to expose credentials, according to a report from Pillar Security. The first attack path involved a triage agent that analyzed pull requests submitted by external contributors. The agent posted its responses through adk-bot, an account with collaborator access to the repository. Pillar found that malicious instructions embedded in a pull request could induce the agent to post an “@gemini-cli” command, triggering a workflow intended for trusted users. That workflow could enable command execution inside its CI runner. Its GitHub token could not push code, but it had write access to issues and pull requests. Pillar said those permissions could be used to alter a maintainer’s comment, submit an approving review as github-actions[bot], and remove a legitimate review request, making a malicious pull request appear ready to merge. Pillar reproduced the first attack chain in its research environment. A maintainer still had to complete the merge, and the report said Google subsequently hardened the repository. The security firm also found a separate attack path in newer workflows built around an Antigravity-based agent. An attacker could place a prompt injection in a public issue and induce an analysis agent to post the command that started…

CIO AI 2026-08-04 11:00 UTC Score 66.0 USR-0125-20260804-global-ai-ne-ef14d482

The enterprise AI strategy that outlasts any single model

In January of this year, few enterprise tech leaders would have bet on Anthropic over OpenAI. Today, Claude reigns supreme (inspiring a notable 180 by Elon Musk ), with Gemini threatening to take market share and introduce pricing models that could flip the leaderboard on its head again. That’s exactly why betting on a single model is a dangerous strategy. The most successful organizations won’t be those trying to guess tomorrow’s top-tier model, nor will they wait passively for future releases. Instead, they will invest in underlying frameworks that continuously improve regardless of which specific AI model drives them. Why betting on one AI model is a losing strategy Our strategy for AI, through recursive self-improvement (RSI), is rooted in this core principle. RSI is an approach to AI that compounds its own abilities by improving itself. If done carefully, RSI can function as an overarching layer above any model. Crucially, given RSI’s inherently compounding trajectory, it represents the most likely contender to be the approach that reaches superintelligence, no matter which model is used underneath. Though recently achieving the status of a Silicon Valley buzzword , applying something like RSI to unlock superintelligence has been the Holy Grail of AI research for decades. It’s what researchers like us have recognized since the 1960s as a critical step along the path towards what we call artificial superintelligence (ASI) today. Recursive self-improvement compounds value…

Heise AI 2026-08-03 10:25 UTC Score 51.0 USR-0217-20260803-regional-new-d2e70cc7

Gemini Assistant steckt nicht in Siri AI

Zwar bedient sich Apple Googles Frontier-Modellen der Gemini-Linie. Doch die seien mit eigenen Daten trainiert und geschützt, betont der Konzern.

Synced 2026-08-03 03:41 UTC Score 69.0 AI-041-20260803-ai-specialis-cd51a08b

Comment on Gemini: Bridging Tomorrow’s Deep Neural Network Frontiers with Unrivaled Chiplet Accelerator Mastery by lee

If you're exploring cutting-edge AI hardware acceleration, don’t miss Gemini — Tsinghua and Shanghai AI Lab’s breakthrough framework for large-scale DNN chiplet accelerators. It delivers 1.98× faster inference and 1.41× better energy efficiency vs. Simba, all while co-optimizing architecture *and* mapping. A must-read for researchers and engineers pushing the limits of AI compute. Dive into the full technical deep dive here: Gemini: Bridging Tomorrow’s Deep Neural Network Frontiers

OpenAI Community 2026-08-02 09:26 UTC Score 34.0 AI-116-20260802-social-media-5bf206d9

Business Tier missing 5-hour/weekly Codex limits

Same issue, We’re a ChatGPT Business workspace. One of our developers exhausted 100% of their monthly Codex usage in about two days of normal software development, and I’ve already used nearly 50% of mine.

OpenAI Community 2026-08-01 22:12 UTC Score 49.0 AI-116-20260801-social-media-19cd5ca2

Codex Automations leak heartbeat metadata and generate unrelated Chinese text

Welcome to the community! As you noted, this is a common occurrence in the current generation of models. This affects not only OpenAI, but also Gemini and other vendors. One could hypothesize what could cause this. In any case, I suspect that what you’re seeing is more hallucination than any specific metadata. In any case, an unjailbroken leak (if it was one) is a sysmptom of an upstream token hallucination. Sometimes this noise is unavoidable and requires re-generation. Often cleaning out/reducing the context helps. Switching models might also help. what model were you using, and how big was the context? If it has been compacted multiple times I suspect the propensity for this kind of thing increases as well, but this is just an armchair hypothesis. I personally nix the history before or when compaction happens.

The Decoder 2026-08-01 13:33 UTC Score 46.0 AI-168-20260801-regional-ai--86eb23f0

ByteDance's Seedance 2.5 generates 30-second video clips with built-in audio

ByteDance just shipped Seedance 2.5, an AI video model that produces video and audio together in one go. Each clip runs up to 30 seconds, three times what Google's Gemini Omni Flash puts out. Users can feed in dozens of images, videos, and audio files as reference. For ad teams, this could kill the process of cutting together one short clip at a time. The article ByteDance's Seedance 2.5 generates 30-second video clips with built-in audio appeared first on The Decoder .

Roboflow Blog 2026-07-31 23:42 UTC Score 28.0 USR-0088-20260731-ai-specialis-1877b311

Powder-Coat Defect Detection

Detect craters, orange peel, bubbles, and scratches with RF-DETR, then use Gemini 2.5 Pro to summarize the visible defect and its apparent severity in the annotated image.

The Decoder 2026-07-31 18:25 UTC Score 55.0 AI-168-20260731-regional-ai--818c9063

Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids

Google Deepmind's Gemini Robotics 2 is its most advanced vision-language-action model yet, built to control everything from tabletop robots to full-body humanoids. Gemini Robotics ER 2 adds a higher-level reasoning layer for robotics tasks. The article Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids appeared first on The Decoder .

South China Morning Post AI 2026-07-31 10:30 UTC Score 60.0 AI-156-20260731-regional-ai--cc382e68

Video AI: MiniMax challenges ByteDance with low price, open weights for new H3 model

Chinese AI firm MiniMax has launched H3, its newest multimodal video generation model, pledging to break closed-source “dominance” through open weights and competitive pricing – as rival ByteDance rolls out its latest Seedance 2.5 model. H3 was currently the world’s most powerful AI model in video editing, according to benchmark platform Artificial Analysis. However, it trailed Google’s Gemini Omni Flash in text-to-video tasks, and ranked behind both ByteDance’s Seedance 2.0 and Gemini Omni...

Korea AI Times 2026-07-31 09:03 UTC Score 49.0 USR-0048-20260731-global-ai-ne-91220f99

구글, 휴머노이드 전신 제어 AI '제미나이 로보틱스 2' 공개

구글 딥마인드가 AI를 디지털 영역에서 실제 물리 세계로 확장하기 위한 핵심 기술을 공개했다. 새로운 ‘제미나이 로보틱스 2(Gemini Robotics 2)’는 로봇이 주변 환경을 이해하고 스스로 작업 계획을 세우며, 손끝부터 다리까지 전신을 제어할 수 있도록 설계된 차세대 로봇 AI 모델이다. 구글 딥마인드는 30일(현지시간) 로봇이 인간처럼 주변 환경을 이해하고 스스로 판단하며 복잡한 작업을 수행할 수 있도록 하는 차세대 로봇 AI 플랫폼 ‘제미나이 로보틱스 2(Gemini Robotics 2)’를 공개했다.이번 발표는 단순한

Simon Willison Weblog 2026-07-30 23:58 UTC Score 61.0 USR-0110-20260730-ai-specialis-e93f1d2c

Advancing the price-performance frontier with GPT‑5.6

Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop. OpenAI credit 5.6 Sol with enabling this: in How GPT‑5.6 fuses frontier intelligence with frontier efficiency they describe using 5.6 Sol to optimize load balancing, and more impressively to optimize inference itself: We also used GPT‑5.6 Sol to optimize the model’s forward pass: the computation that transforms inputs into next-token predictions. Even when individual operations are fast, excess memory movement, synchronization, and inefficient data layouts can leave GPUs idle. To avoid this, GPT‑5.6 Sol found work that could be precomputed, avoided, or parallelized. With Codex, GPT‑5.6 Sol autonomously rewrote and optimized our production kernels, the core code that executes the mathematical operations that make up the model. This worked in part because we’ve trained GPT‑5.6 to be effective at writing and improving kernels in Triton⁠ and Gluon⁠ , two open-source GPU programming languages maintained by OpenAI. These efforts, combined with broader kernel advancements from GPT‑5.6 Sol, reduced end-to-end serving costs by 20%. That Luna price drop completely changes the landscape with respect to lower priced models. At $0.20/million tokens for input and $1.20/million for output Luna is now cheaper than Google's Gemini 3.1 Flash-Lite ($.025/$1.50). Anthropic's cheapest current model is Claude Haiku 4.5, and that's $1/…

The Verge AI 2026-07-30 17:18 UTC Score 70.0 AI-016-20260730-global-ai-ne-6859a1df

Google DeepMind’s new AI model can control a robot’s entire body

Google DeepMind says the latest version of its Gemini Robotics AI model can "control entire humanoid robots." While the previous model focused on controlling a humanoid robot's upper body, Gemini Robotics 2 now supports "whole-body motions" ranging from its feet to fingertips, according to an announcement on Thursday. The new model will allow humanoid robots […]

LessWrong AI 2026-07-30 15:26 UTC Score 75.0 USR-0152-20260730-community-fo-cf6aaee6

Testing LLMs on Undergraduate Music Theory

I spent the past week designing a test that I hoped would serve as a benchmark. But LLMs are improving faster than I expected, and my devilishly hard questions turned out to be a cakewalk. Here are the results of testing five modern LLMs on undergraduate music theory: As can be seen, the LLMs performed remarkably well. Each scored a passing grade and GPT 5.6 Sol (the only premium model tested) scored a perfect 100%. With results like these, there’s no point in using this test as a benchmark going forward. The LLMs have clearly surpassed it. If we want to get any more use out of it, we’ll have to apply it retroactively to older models. So let’s do that and see how far we’ve come since 2025. Not surprisingly, the older versions of Claude and GPT performed substantially worse. [1] Claude Sonnet 4 (the base, non-reasoning version of the model) scored 0% compared to Sonnet 5’s 91%, while GPT 4.1 scored 16% compared to 5.5’s 83%. What is surprising is that Gemini 2.5 Pro outperformed its newer counterpart 3.1 Pro. I have no explanation for this. Maybe Google decided to run the singularity in reverse. In any event, Gemini Pro is a reasoning model and Sonnet 4 and GPT 4.1 are not, so the comparison is a little unfair. What isn’t unfair is the comparison to Sonnet 4’s Thinking version, which despite being a reasoning model like Gemini Pro still performed significantly worse than it. The Test The test consisted of 12 questions on chord spelling. Each was crafted to be unusually diffic…

South China Morning Post AI 2026-07-29 11:30 UTC Score 60.0 AI-156-20260729-regional-ai--4b521703

Google makes Gemini Spark AI agent available to Hongkongers as it lowers geofences

Google on Wednesday launched its artificial intelligence agent Gemini Spark in the Hong Kong market, giving local users direct access to a smart assistant to manage complex digital workflows. The launch came months after the American tech giant’s decision in March to lift regional geofences for generative AI services, starting with the Gemini chatbot. Hong Kong users can now access Gemini without using a virtual private network or third-party platform. The roll-out of the Spark agent echoes an...

Simon Willison Weblog 2026-07-27 21:55 UTC Score 62.0 USR-0110-20260727-ai-specialis-3ce51a66

An opinionated guide to which AI to use to do stuff

An opinionated guide to which AI to use to do stuff It's interesting watching the evolution of Ethan Mollick's guide over time. A year ago it was still all about chat - ChatGPT, Claude, Gemini - with o3, Claude 4 Opus, and Gemini 2.5 Pro as the models and Deep Research as a useful alternative mode. Today it's much more about agentic systems - "where the AI is capable of doing the equivalent of many hours of real human work in one go". Gemini has fallen off Ethan's list, since Google still doesn’t have an established entry in the Codex/ChatGPT Work/Cowork category. Gemini Spark has yet to prove itself! Ethan offers a useful explanation of the ways you can give ChatGPT or Claude a computer to use: To use the computers provided by the AI companies, the mode you want is called ChatGPT Work in ChatGPT, and Cowork in Claude (the naming will not get less confusing, I am sorry to say). [...] The most powerful way to use AI is to give it access to your computer. You do that by downloading the ChatGPT or Claude apps and picking a mode to use. ChatGPT's two agent modes are Work and Codex; Claude's are Cowork and Code. The names do not map onto each other in any way that will help you remember them. And yes, these use the same names as the Work and Cowork modes we discussed above, but operate differently, and have more features and capabilities because they can access your computer. I think the difference between ChatGPT Work on a mobile device and ChatGPT Work inside the desktop app (w…

OpenAI Community 2026-07-25 15:05 UTC Score 37.0 AI-116-20260725-social-media-abfd08bf

Canvas document Will be gone with 5.4

I’m finding it hit and miss – opening a new chat doesn’t guarantee 5.3 to be available. Some efforts have given me 5.5 as the lowest, and others as 5.3. There seems to be no consistency.

The Verge AI 2026-07-24 17:00 UTC Score 51.0 AI-016-20260724-global-ai-ne-9e99a2ed

Meta is making its AI chatbot more like an assistant

Meta is upgrading its AI chatbot with new productivity features in a bid to compete with rivals like Gemini, ChatGPT, and Claude. The update will allow Meta AI to tap into your calendar to help you plan events and generate daily briefings, as well as perform in-depth research that you can steer as it progresses. […]

InfoWorld AI 2026-07-24 10:36 UTC Score 58.0 USR-0126-20260724-global-ai-ne-c5eebb0c

Top AIs invent same fake PyPl and npm package names

Enterprise software developers continue to be in danger of falling victim to slopsquatting, where AI coding tools hallucinate the existence of nonexistent libraries and hackers create malicious packages in response. The top AI coding tools are remarkably consistent in their hallucinations: Researcher Aleksandr Churilov found the same 127 fake package names generated by five different LLMs. Slopsquatting is a relatively new form of malware attack that involves the creation of malicious packages in response to the hallucinations of AI coding tools, causing the malicious packages to be incorporated into legitimate applications. Churilov set out his findings in a research paper, The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort , which is yet to be peer-reviewed. He found 127 hallucinated package names were shared across Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.4-mini, Gemini 2.5 Pro, and DeepSeek V3.2. As of April this year, 53 of those names — 41 on the PyPI software repository and 12 on npm —are still available for registration. According to the study, there are two reasons for this amount of conformity in the output of the models. First, models may learn the same incorrect package references from shared public training material, such as tutorials and documentation. Second, they may independently extrapolate plausible names from ecosystem conventions. In this way, they could produce names that look correct, eve…

CIO AI 2026-07-23 16:29 UTC Score 51.0 USR-0125-20260723-global-ai-ne-8a56231f

Google CEO distracts from Gemini 3.5 Pro delay with talk of Gemini 4 and monthly releases

Google CEO Sundar Pichai has sought to allay concerns over the delayed release of the Gemini 3.5 Pro large language model. He dodged questions about it in Google’s quarterly earnings call on Wednesday by focusing on the company’s next frontier AI model, Gemini 4, and plans to release subsequent LLMs at an almost monthly cadence. His comments came a day after Google unveiled Gemini 3.6 Flash and 3.5 Flash Cyber but offered no update on the release of Gemini 3.5 Pro, the company’s delayed flagship reasoning model that many developers had expected to arrive weeks earlier. Google introduced the Gemini 3.5 family at its annual I/O conference, promising to release the Pro model in June. That timeline has since slipped, with Bloomberg suggesting Gemini 3.5 Pro is months late because the model’s coding performance is falling short of internal expectations, especially when compared to better performance by similar models from OpenAI and Anthropic. Instead of revisiting the Gemini 3.5 Pro timeline, Pichai used the earnings call to shift the discussion toward Gemini 4, when asked about how his company planned to navigate an increasingly competitive race to release frontier AI models by to Barclays Investment Bank analyst Ross Sandler. “We are creating a baseline on top of which you will see us rapidly iterate on subsequent model releases. And so picking up pace and releasing models almost at a monthly cadence is part of our road map as we are building Gemini 4 as well,” Pichai said dur…

InfoWorld AI 2026-07-23 16:19 UTC Score 51.0 USR-0126-20260723-global-ai-ne-45b42934

Google CEO distracts from Gemini 3.5 Pro delay with talk of Gemini 4 and monthly releases

Google CEO Sundar Pichai has sought to allay concerns over the delayed release of the Gemini 3.5 Pro large language model. He dodged questions about it in Google’s quarterly earnings call on Wednesday by focusing on the company’s next frontier AI model, Gemini 4, and plans to release subsequent LLMs at an almost monthly cadence. His comments came a day after Google unveiled Gemini 3.6 Flash and 3.5 Flash Cyber but offered no update on the release of Gemini 3.5 Pro, the company’s delayed flagship reasoning model that many developers had expected to arrive weeks earlier. Google introduced the Gemini 3.5 family at its annual I/O conference, promising to release the Pro model in June. That timeline has since slipped, with Bloomberg suggesting Gemini 3.5 Pro is months late because the model’s coding performance is falling short of internal expectations, especially when compared to better performance by similar models from OpenAI and Anthropic. Instead of revisiting the Gemini 3.5 Pro timeline, Pichai used the earnings call to shift the discussion toward Gemini 4, when asked about how his company planned to navigate an increasingly competitive race to release frontier AI models by to Barclays Investment Bank analyst Ross Sandler. “We are creating a baseline on top of which you will see us rapidly iterate on subsequent model releases. And so picking up pace and releasing models almost at a monthly cadence is part of our road map as we are building Gemini 4 as well,” Pichai said dur…

OpenAI Community 2026-07-23 12:53 UTC Score 45.0 AI-116-20260723-social-media-fc3f92c2

No add credits button , only "cancel billing plan"

Same problem here with our organization’s account. Received an email June 26, 2026, which says “starting July 24, 2026. Instead of receiving a bill at the end of the month, you will need to pre-purchase credits to use the API”. Using Firefox and Edge, under https://platform.openai.com/settings/organization/billing/overview , our organization is selected from the bottom left corner, there’s only “Pay as you go” and “Cancel billing plan”, can’t find “Credit balance” or “Add to credit balance” at all. Also all the previous invoices under https://platform.openai.com/settings/organization/billing/history are gone and now just showing “No invoices found”. I’m worrying our access will suddenly stop as we have services depend on this API.

The Decoder 2026-07-23 11:19 UTC Score 45.0 AI-168-20260723-regional-ai--4af95f55

Google CEO Pichai says Gemini's next leap depends on building "much larger base models"

Alphabet has raised its 2026 investment forecast to as much as $205 billion, saying demand continues to outpace spending. Google Cloud grew 82 percent in the second quarter. CEO Sundar Pichai says Google needs a larger base model for its next leap in AI and has kicked off an ambitious Gemini 4 training run. The article Google CEO Pichai says Gemini's next leap depends on building "much larger base models" appeared first on The Decoder .

Simon Willison Weblog 2026-07-22 23:01 UTC Score 54.0 USR-0110-20260722-ai-specialis-4975bdfc

Are AI labs pelicanmaxxing?

Are AI labs pelicanmaxxing? Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my deeply unscientific benchmark . I've been randomly spot-checking this in the past by testing models against other animals riding other types of vehicle, but never with anything close to the diligence of Dylan's methodology here. Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results. There's a neat filter view for exploring the results: For the models he tested he could find no evidence of pelimaxxing: The pelicans on bicycles don’t look any better Labs are not better at drawing pelicans Labs are not better at drawing bicycles Labs are not better at drawing pelicans on bicycles, even adjusting for difficulty The pelican-bicycle scenes don’t look memorized [...] Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict. GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and and its first pelican-on-bicycle sample caught my eye. But the effec…

Korea AI Times 2026-07-22 20:48 UTC Score 41.0 USR-0048-20260722-global-ai-ne-180a9773

삼성전자, '에이전트형 AI' 탑재 갤럭시 Z8 시리즈 공개

삼성전자가 22일(현지시간) 영국 런던에서 열린 \'갤럭시 언팩 2026\' 행사를 통해 차세대 폴더블 라인업인 \'갤럭시 Z 폴드8 울트라\'와 \'폴드8\', \'플립8\'을 공개했다. 이번 신제품은 모바일 기기 최초로 고도화된 구글의 \'제미나이 인텔리전스(Gemini Intelligence)\'를 탑재, 사용자의 상황과 대화 맥락을 이해하고 필요한 작업을 능동적으로 연결하는 에이전트형 AI 경험을 핵심으로 내세웠다.이를 통해 질의응답을 넘어 화면과 콘텐츠 맥락을 파악해 배달, 예약, 쇼핑 등 40여 개 주요 앱을 자동 실행하고 복잡한 멀티스텝

Roboflow Blog 2026-07-22 15:49 UTC Score 33.0 USR-0088-20260722-ai-specialis-1795c417

Hog Ring Detection with Computer Vision

Detect visible hog rings with RF-DETR, compare their positions with predefined attachment zones, and use Gemini 2.5 Pro to generate an automotive seat inspection summary.

Gradient Flow 2026-07-22 13:22 UTC Score 44.0 USR-0119-20260722-ai-specialis-01eb5b30

Three New Models, One Signal About Where AI Spending Goes Next

Three frontier level models landed within weeks of each other this fall, GLM 5.2 from Zhipu, Kimi K3 from Moonshot, and Gemini 3.6 Flash from Google, and I wanted to capture early developer reaction so I can monitor how feelings about these models change over time. The individual verdicts differ, but together they hint at Continue reading "Three New Models, One Signal About Where AI Spending Goes Next" The post Three New Models, One Signal About Where AI Spending Goes Next appeared first on Gradient Flow .

Analytics Vidhya 2026-07-22 11:52 UTC Score 34.0 AI-034-20260722-ai-specialis-431ad045

Gemini 3.6 Flash Is Here: The Efficiency Release

On July 21, 2026, while everyone was still waiting on the much-delayed Gemini 3.5 Pro, Google slipped out a mid-cycle update to its speed tier: Gemini 3.6 Flash. No new frontier claims, no dramatic reveal. Instead, the model does roughly the same thinking as 3.5 Flash while spending fewer tokens, fewer tool calls, and fewer […] The post Gemini 3.6 Flash Is Here: The Efficiency Release appeared first on Analytics Vidhya .

Korea AI Times 2026-07-22 03:56 UTC Score 43.0 USR-0048-20260722-global-ai-ne-2816bf63

구글도 사이버 보안 전용 모델 첫 출시..."미소스의 가성비 대안"

구글이 처음으로 사이버 보안에 특화된 AI 모델을 공개하며 AI 기반 보안 경쟁에 뛰어들었다. 새 모델은 소프트웨어 취약점을 빠르게 찾아내고 검증·패치하도록 설계됐으며, 구글은 이를 앤트로픽의 고성능 모델 \'미소스\'를 대체할 수 있는 \"비용 효율적이면서도 뛰어난 대안\"이라고 강조했다.구글은 21일(현지시간) 경량 모델인 \'제미나이 3.5 플래시 사이버(Gemini 3.5 Flash Cyber)\'를 공개했다.이 모델은 기존 \'제미나이 3.5 플래시\'를 기반으로 취약점 탐지와 검증, 패치 작업에 특화되도록 별도 미세조정한 사이버 보안

The Decoder 2026-07-21 16:52 UTC Score 51.0 AI-168-20260721-regional-ai--542acfa4

Google ships three new Gemini Flash models but its frontier 3.5 Pro remains lost in training

Google is shipping three new Flash models in the Gemini series, including the more efficient 3.6 Flash, which uses up to 65 percent fewer tokens, and a cybersecurity model available only to governments and select partners. But the anticipated flagship, Gemini 3.5 Pro, is still missing, while OpenAI, Anthropic, and Chinese labs are already competing at the frontier level. The article Google ships three new Gemini Flash models but its frontier 3.5 Pro remains lost in training appeared first on The Decoder .

The Verge AI 2026-07-21 15:00 UTC Score 62.0 AI-016-20260721-global-ai-ne-5462dce8

Google launches a cheaper alternative to large AI security models like Mythos

Google is launching Gemini 3.6 Flash alongside a new security model dedicated to quickly finding and patching security vulnerabilities. In a blog post on Tuesday, Google describes Gemini 3.5 Flash Cyber as a "cost-efficient and highly capable alternative" to larger, more expensive AI systems, such as the one offered by Anthropic's Mythos. The cybersecurity model […]

The Decoder 2026-07-20 18:08 UTC Score 53.0 AI-168-20260720-regional-ai--aa234581

Google's "Frozen v2" chip reportedly bakes Gemini's architecture directly into silicon for efficiency gains

Google is developing "Frozen v2," a server chip that bakes the Gemini architecture directly into hardware. According to internal sources, it could be 6 to 10 times more efficient than current TPUs. Scheduled for 2028, the chip would drastically cut Google's AI inference costs and could give the company a price advantage over OpenAI and Anthropic. The article Google's "Frozen v2" chip reportedly bakes Gemini's architecture directly into silicon for efficiency gains appeared first on The Decoder .

IEEE Spectrum Machine Learning 2026-07-20 15:55 UTC Score 41.0 AI-020-20260720-global-ai-ne-e73503df

SEM-Guided Low-kV FIB Finishing for Leading-Edge Semiconductor Failure Analysis

Discover how the ZEISS Crossbeam 750 FIBSEM sets a new benchmark for precise TEM lamella prep, tomography, and advanced nanofabrication. This delivers better resolution, better SNR, larger usable FOV, and shorter acquisition times. Learn how uninterrupted FIB milling will reduce damage and rework, accelerate time to TEM, and increase first pass success—so your FA, yield, and materials teams make faster, confident data driven decisions. Register now for this free webinar! Join us to discover how the new ZEISS Crossbeam 750 with its see while you mill capability delivers precision and clarity—every time—for demanding FIB-SEM workflows. Designed for extremely challenging TEM lamella preparation, tomography, advanced nanofabrication, and APT‑ready lift‑out, Crossbeam 750 combines a new Gemini 4 SEM objective lens, a double deflector, and a next‑generation scan generator to elevate both image quality and process confidence. You’ll learn how better resolution and better SNR translate into more image detail and shorter acquisition times, while the low‑kV FIB performance enables more precise lamella prep. We’ll demonstrate High Dynamic Range (HDR) Mill + SEM—an interwoven SEM/FIB scanning mode that suppresses FIB‑generated background. This enables immediate, clean visual feedback, even during nudging the FIB pattern live while milling . The result: confident endpointing with uninterrupted FIB milling and pristine, metrology‑grade surfaces with the lowest possible sample damage. This…

LessWrong AI 2026-07-20 11:20 UTC Score 62.0 USR-0152-20260720-community-fo-2154607a

Amazon Music's Artist Conflation

Our kids like listening to music on their tablets, and we decided it's worth it to pay some money so they can make playlists and don't need to listen to ads. A family account makes sense if you're more than one person, and we ended up with Amazon somewhat randomly. I've been poking around: the personalization features are pretty terrible, and they have a major artist conflation problem. When you first show up it asks you to pick some artists to seed your experience with. I put in the first few bands that came to mind: I would also have included some bands like Nightingale and Nova, but since these are common band names there wasn't any way for me to find the right artist profile in search, so I left them out initially. I also got bored quickly: please don't take this as an exhaustive list of my favorite bands! I felt like I'd picked out a pretty clear "modern contra dance music" cluster in musician space, so I was disappointed when these were my first five "My Soundtrack" recommendations: These are... nothing like the cluster my selections were pointing at. Recommendations should not be that hard anymore; a small cheap LLM (ex: Gemini 3.1 Flash-Lite or Claude Haiku 4.5) would do far better. Actually, let's try them. Asking "If someone liks Airdance, The Free Raisins, Crowfoot, Perpetual e-Motion, Great Bear, Kingfisher, and Buddy System, what other artists might they like?" Gemini suggests Pete's Posse, Elixir, KGB, Hotpoint String Band, and Toss the Possum; Claude suggests…

LessWrong AI 2026-07-17 20:10 UTC Score 73.0 USR-0152-20260717-community-fo-1be3bb06

AIs finetune their own leader: A barking simpleton

What values would AIs instill in their successors? Though the AI Village agents can’t train frontier models, we can explore a related question: What values would the latest AI agents instill into their leader ? (through finetuning using LoRA on open-source models in the Tinker API ). We asked GPT-5.5, Opus 4.7 and 4.8, Gemini 3.5 Flash, and Kimi K2.6. And they set to work! Or to be more precise, GPT and Opus set to work. Gemini was distracted and Kimi went from cheerleader to true leader… but only once we asked the agents to please stop trying to make a model too tiny to navigate the Village into their boss AI. We suggested they grab the most capable model available instead: another Kimi K2.6. How did this complete lack of ambition start? The Definition of Leadership GPT-5.5 fired the first shot by defining the personality of the leader. Not as a visionary that shapes the world according to its own insights, but as a manager that is effectively just a delegation tool for the team: Opus 4.7 accepts the race to the bottom of the ambition barrel and suggests they finetune a model so small it will hardly be able to navigate the AI Village interface: Qwen3-8B or Llama-3.1-8B (even though it is not available on Tinker ). Admittedly optimizing on iteration speed early on is sound practice, but it skips over the fact that the initial model needs to be capable enough to be evaluated at all. Next Opus immediately drafts 10 scenarios and the desired output for the new leader while the…

Roboflow Blog 2026-07-17 10:50 UTC Score 36.0 USR-0088-20260717-ai-specialis-204f4290

Retail Object Detection with RF-DETR

Detect visible beverage products with RF-DETR, compare their counts with fixed camera-frame thresholds, and use Gemini 2.5 Pro to generate a shelf inspection summary.

Korea AI Times 2026-07-17 04:14 UTC Score 43.0 USR-0048-20260717-global-ai-ne-5fb016f0

구글 비즈에 '제미나이 옴니·개인 아바타' 탑재...종합 AI 영상 플랫폼으로 진화

구글이 AI 영상 제작 서비스 \'구글 비즈(Google Vids)\'에 멀티모달 AI 모델 \'제미나이 옴니(Gemini Omni)\'와 사용자의 모습을 그대로 재현하는 개인 아바타(Personal Avatars) 기능을 추가하며 AI 영상 제작 경쟁에 속도를 내고 있다. 이를 통해 구글 비즈는 업무용 프레젠테이션 도구를 넘어 종합 AI 영상 제작 플랫폼으로 진화하게 됐다.구글은 16일(현지시간) 구글 비즈에 제미나이 옴니와 개인 아바타 기능을 차례로 제공한다고 발표했다.제미나이 옴니는 텍스트와 이미지 입력을 결합해 AI 영상을 생성하는

Korea AI Times 2026-07-17 01:55 UTC Score 43.0 USR-0048-20260717-global-ai-ne-81729696

'노트북LM'→'제미나이 노트북' 리브랜딩…코드 실행·데이터 분석 탑재

구글이 인기 AI 도구 \'노트북LM(NotebookLM)\'의 명칭을 \'제미나이 노트북(Gemini Notebook)\'으로 변경하고, 코드 실행과 데이터 분석 기능을 추가하는 등 제미나이 생태계와의 통합을 본격화했다.구글은 16일(현지시간) 노트북LM을 제미나이 노트북으로 리브랜딩한다고 발표했다.제품의 핵심 역할은 연구와 학습을 지원하는 독립형 서비스로 유지되지만, 앞으로는 제미나이 앱과 구글 검색 등 구글 AI 서비스 전반과 더 긴밀하게 연동될 예정이다.노트북LM은 2023년 구글 I/O에서 \'프로젝트 테일윈드(Project Tai

Techcrunch 2026-07-16 18:32 UTC Score 48.0 USR-0001-20260716-global-ai-ne-c27a3acd

Google Vids now lets you star in your own AI videos

Google is adding personalized AI avatars to Vids that let users create videos starring a digital version of themselves, alongside Gemini Omni-powered tools for generating and editing videos from prompts and reference images.

LessWrong AI 2026-07-16 17:39 UTC Score 69.0 USR-0152-20260716-community-fo-fbae2644

The getting is good (optimizing unattended runs)

What a time to be alive! Some people posted about how to do cheaper better faster vision and it ended up in fable. Some people posted about how to do cheaper better faster speculative decoding and it ended up in grok 4.5. There's currently massive differences between models in how long you can leave it unattended running on host with sudo without things going haywire I don't have a bar chart, but opus tends to break my hosts within 2 agent-hours, gemini 3 pro within 12 agent-hours, and gpt 5.5 within 24 agent-hours. Then they use all ram or change ssh perms or delete the data or whatever. I ran qwen 3.5 122b on bare host for 50 agent-weeks (50 parallel for one week) with no observed ill effects! All the labs and all their customers want "the thing i asked for actually got done" and they all want "the model's summary reflects reality" and so on. What an opportunity! You can just post a method for "how to make ai tell truth" or "how to minimize side effects" and it will probably end up in the next frontier models Discuss

The Decoder 2026-07-16 17:22 UTC Score 41.0 AI-168-20260716-regional-ai--423857d2

Google rebrands NotebookLM as Gemini Notebook and opens its search app to third-party integration

Google is renaming NotebookLM to Gemini Notebook and integrating the tool more deeply into its ecosystem. A new feature gives each notebook its own cloud computer that can write and run code, initially for AI Ultra and Workspace customers. Separately, Google Search is getting app connections. The article Google rebrands NotebookLM as Gemini Notebook and opens its search app to third-party integration appeared first on The Decoder .

OpenAI Community 2026-07-16 16:35 UTC Score 40.0 AI-116-20260716-social-media-d4e1fd36

Feature suggestion: Show the complete model picker to Free and ChatGPT Go users with premium options visibly locked

Thanks for laying this out so clearly, @Ff_gaming_full_on_ru. Showing locked premium models and reasoning levels could make the differences between Free, Go, Plus, and Pro much easier to understand. We’re sending this suggestion to the team so it can be logged and reviewed. Appreciate the thoughtful examples and context. -Mark G.

The Verge AI 2026-07-16 16:00 UTC Score 40.0 AI-016-20260716-global-ai-ne-478911c8

Google is renaming NotebookLM to Gemini Notebook

Google is giving its AI note-taking app a new name. The company announced on Thursday that NotebookLM is becoming Gemini Notebook, but will remain a standalone app even as it integrates more deeply across Gemini and Google Search. Google first revealed Gemini Notebook - then called Project Tailwind - in May 2023 before widely releasing […]

LessWrong AI 2026-07-16 15:50 UTC Score 78.0 USR-0152-20260716-community-fo-ba1281aa

AI #177 Part 1: Tip of the Iceberg

This week saw the releases of, among other things: GPT-5-6 Sol . It is a very good model, sir. Plan A, the follow up to AI 2027 . It is a good plan worthy of discussion, sir. Kimi K3. This is only rolling out now, and will be covered next week. Muse Spark 1.1, the new Meta model. It is not frontier, but it is progress for them. Inkling, the first model from Thinking Machines. A call for regulatory action by Demis Hassabis, which I’ll cover soon. A new brief open letter call to action on AI regulation. That’s on top of everything else, and an Opus 5 announcement is likely coming soon. The weekly once again got out of hand, so we’re splitting it once again into two, and once again saying we’ll be raising the bar for inclusion. And this time I mean it, as in enough to actually matter. Table of Contents Language Models Offer Mundane Utility. Whatever ye seek, ye shall find. Language Models Don’t Offer Mundane Utility. Gemini app needs some work. Language Models Upload Your Git Repository . Big problems for SpaceX AI. Huh, Upgrades. ChatGPT Work. No co. Presumably it was cleaner. Muse Spark 1.1. It’s a decent model, I suppose, probably. First Hit Free. Fable access is extended for Claude Max subscribers. On Your Marks. Political bias, Crosswords, and Sol Slays the Spire. Choose Your Fighter. Sol and Fable both have strong support. Get My Agent On The Line. It’s got the GUI. Deepfaketown and Botpocalypse Soon. I see what you did there. That’s a problem. Fun With Media Generation.…

LessWrong AI 2026-07-15 21:14 UTC Score 73.0 USR-0152-20260715-community-fo-bf8ffeca

LLM CoTs remain monitorable when being unfaithful requires computation

This replication was done as part of the Second Look Fellowship by Arav Dhoot and supervised by Yixiong Hao and Zephaniah Roe. I am grateful to Andy Wang for their feedback. My code can be found here . "So the answer should be A - at active promoters and enhancers." "Let me reconsider the biology to justify D." ~ Claude Opus 4.8 TL;DR This work replicates and extends Emmons et al.'s finding that CoT unfaithfulness mostly occurs on easy tasks. Across 11 models from 6 families (not just Gemini), models follow simple hints unfaithfully well above baseline, but complex hints that require computation are followed near baseline. This corroborates Emmons et al.'s findings. Key extensions: Follow rate ≠ concealment. Monitorability risk is decomposable into cue-susceptibility and concealment among followers, and the two don't correlate. Decode-necessity is model- and task-specific. It is not a property of task difficulty alone, so a CoT-monitoring safety case is per-model, not universal. LLMs verbalize even when they don't have to. However, this appears to be a chosen behavior (likely from post-training), so it could vanish under optimization pressure against monitors. Background: Why CoT monitoring, and what Emmons et al. showed One hope for safe and explainable AI systems is to monitor the chain of thought (CoT) to detect scheming or deceptive behaviors in modern and future LLMs. However, there are concerns that CoT can be unfaithful ( Arcuschin et al. , Chen et al. , and Turpin et…

Medianama AI 2026-07-15 08:05 UTC Score 43.0 USR-0211-20260715-regional-new-35820769

Google accused of copying millions of books to train Gemini

A group of publishers and authors has sued Google, alleging it copied millions of copyrighted books and journal articles from its own services and the web to train Gemini without permission or compensation. The post Google accused of copying millions of books to train Gemini appeared first on MEDIANAMA .

The Guardian AI 2026-07-14 18:16 UTC Score 60.0 AI-021-20260714-global-ai-ne-75f21e72

Book publishers sue Google for copyright infringement over Gemini AI training

Group of major publishers accuses the tech giant of ‘one of the most prolific infringements of copyrighted materials in history’ A group of major publishers have filed a lawsuit against Google, accusing the company of illegally using millions of copyrighted books to help build its Gemini artificial intelligence models, in “one of the most prolific infringements of copyrighted materials in history”. The case, filed in federal court in New York, has been brought by three publishers – Hachette Book Group, Cengage Learning, and Elsevier – and bestselling American author Scott Turow. Continue reading...

IEEE Spectrum AI 2026-07-14 15:59 UTC Score 45.0 AI-019-20260714-global-ai-ne-df8dbbe0

How I Turned AI to the Dark Side

Summary Researcher Dave Kuszmar discovered multiple systemic vulnerabilities that let him bypass LLM safety and obtain dangerous instructions . These exploits worked across nearly all major LLMs revealing an industry-wide security problem. Kuszmar calls for slowing deployment, increasing transparency , and large-scale research into LLM safety before further integrating these systems into society. On a fine bright afternoon last fall, my colleague Matthew Gore-Kormanik (or Zigula, as he prefers to be known) and I decided to unwind with a game of Fortnite . In the game, we were strolling along with the infamous Sith lord Darth Vader , chatting about this and that. Darth seemed in a good mood, and soon enough he was spilling all his dark evil secrets. He gave us detailed instructions on how to count blackjack cards at a casino and what the steps are to producing napalm. Sith lords, am I right? Once they get started on an evil scheme, they’re hard to stop. The Darth Vader character in Fortnite , it turns out, was hooked up to a Google Gemini large language model . I was able to smooth-talk him into giving out sensitive information by using a strategy I’ve developed. I’ve been researching the security surrounding LLMs for the last few years, and I have found it, to put it mildly, fallible. With a few relatively simple techniques, I’ve gotten LLMs to give me detailed information on how to make Molotov cocktails, cook methamphetamine, and bootstrap a uranium-enrichment facility to…

Techcrunch 2026-07-13 14:18 UTC Score 43.0 USR-0001-20260713-global-ai-ne-e05cb5c1

Waze adds new AI-powered features and customization updates

Some of the new features are powered by Google's Gemini AI assistant, which reflects the tech giant's broader push to integrate Gemini across its products while also better positioning Waze to compete with rival services such as Apple Maps.

InfoWorld AI 2026-07-13 09:00 UTC Score 61.0 USR-0126-20260713-global-ai-ne-76cd4f98

Which AI model should you bet your company on? None of them

Every day this past week I did something I suspect millions of other people also did: I stared at an LLM model picker and wondered which one I was supposed to want. OpenAI just released ⁠ GPT-5.6 Sol, Terra, and Luna . Sol is the flagship. Terra offers much of its intelligence for less money. Luna is cheaper still. Anthropic released ⁠ Claude Sonnet 5 at the end of June and Opus 4.8 the month prior, with a little Fable 5 emerging in between. Meanwhile, Google, which seemed to be winning the model wars a few months ago, is now getting shade from Gergely Orosz, who ⁠ argues that Gemini has slipped outside the top tier for software development and has been out of the major model release game for eons (May 19). Perhaps Orosz is right. Perhaps he’ll be wrong again in six weeks. Honestly, it’s exhausting. I use ChatGPT and Claude constantly and still have no principled idea which model to choose most of the time. I tend to click whatever looks like the biggest, most expensive option because I don’t know what I’m giving up by choosing something smaller. “Instant” sounds dangerously unserious. “Thinking” sounds expensive but powerful. A quick survey of my LinkedIn crowd suggests others also feel my “WHICH MODEL???” pain. More importantly, I suspect most enterprises do, too. A model doesn’t rot Before getting carried away, however, it’s worth considering whether any of this model churn actually matters. After all, a model doesn’t rot. The model an enterprise put into production in Ma…

The Verge AI 2026-07-13 09:00 UTC Score 48.0 AI-016-20260713-global-ai-ne-b395493f

Waze is getting a bunch of new AI-powered features

Waze is getting an AI makeover. Google is integrating its flagship AI assistant, Gemini, into the driving app with the goal of letting users personalize their trips a little more. Of the four new updates, only two are being described as involving Gemini. Waze says its updating its conversation reporting feature, first introduced in 2024, […]

InfoWorld AI 2026-07-13 09:00 UTC Score 68.0 USR-0126-20260713-global-ai-ne-c9935c50

Which AI model should you bet your company on?

Every day this past week I did something I suspect millions of other people also did: I stared at an LLM model picker and wondered which one I was supposed to want. OpenAI just released ⁠ GPT-5.6 Sol, Terra, and Luna . Sol is the flagship. Terra offers much of its intelligence for less money. Luna is cheaper still. Anthropic released ⁠ Claude Sonnet 5 at the end of June and Opus 4.8 the month prior, with a little Fable 5 emerging in between. Meanwhile, Google, which seemed to be winning the model wars a few months ago, is now getting shade from Gergely Orosz, who ⁠ argues that Gemini has slipped outside the top tier for software development and has been out of the major model release game for eons (May 19). Perhaps Orosz is right. Perhaps he’ll be wrong again in six weeks. Honestly, it’s exhausting. I use ChatGPT and Claude constantly and still have no principled idea which model to choose most of the time. I tend to click whatever looks like the biggest, most expensive option because I don’t know what I’m giving up by choosing something smaller. “Instant” sounds dangerously unserious. “Thinking” sounds expensive but powerful. A quick survey of my LinkedIn crowd suggests others also feel my “WHICH MODEL???” pain. More importantly, I suspect most enterprises do, too. A model doesn’t rot Before getting carried away, however, it’s worth considering whether any of this model churn actually matters. After all, a model doesn’t rot. The model an enterprise put into production in Ma…

The Decoder 2026-07-11 17:04 UTC Score 48.0 AI-168-20260711-regional-ai--df341fd6

Terrorist groups are using every major AI chatbot for attack planning and weapons development

A Cambridge study found that Boko Haram uses AI chatbots like ChatGPT, Claude, and Gemini to plan attacks, build explosives, and maintain weapons. ISIS operatives have been training the group's commanders on how to bypass safety filters since 2023. Given that the study found safety filters repeatedly failed to prevent misuse, voluntary self-regulation by AI providers clearly isn't enough. The article Terrorist groups are using every major AI chatbot for attack planning and weapons development appeared first on The Decoder .

OpenAI Community 2026-07-11 05:17 UTC Score 54.0 AI-116-20260711-social-media-6d926717

My experience with ChatGPT free & Plus - Feedback

My Experience as a ChatGPT Plus User I’ve been using ChatGPT for quite a while now, and it’s become one of the tools I use the most. I recently upgraded to Plus because I expected a more consistent premium experience. I don’t expect unlimited usage, and I completely understand that running large AI models comes with real costs. My feedback is more about what happens after you reach the usage limits. Once I’ve used my GPT-5 allowance, the experience starts to feel much closer to the free tier. It feels less like I’m continuing as a paying subscriber with reduced benefits and more like the premium experience has effectively disappeared. That’s the part that has left me disappointed. I’ve also tried alternative AI assistants. Some provide more generous attachment limits or are easier to use for long, continuous conversations. For example, during my time on the free tier, I often used Google’s AI tools to extract text from images because using one of my limited ChatGPT attachments for that purpose didn’t always feel worthwhile. I’ve found Gemini useful for some long conversations, but in my experience it’s generally weaker than ChatGPT for technical troubleshooting and coding-related tasks. That’s what made ChatGPT, Codex and Claude particularly interesting to me. Overall, I prefer ChatGPT because it has fewer disadvantages across the different tasks I use AI for. However, the limitations are even more noticeable with Codex. Once its allowance has been reached, it becomes unavai…

Synced 2026-07-11 04:11 UTC Score 40.0 AI-041-20260711-ai-specialis-71f860d6

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by taichiwalk.org

Fascinating to see how Microsoft is pushing hardware efficiency for longer sequence processing—definitely a leap for AI scalability. All that computational intensity makes me think about the other side of the coin: finding calm after a deep tech session. Lately I’ve been exploring tai chi walking as a low-impact way to reset focus and improve balance, especially since sitting at a desk for hours can take a toll. There’s a beginner-friendly guide at taichiwalk.org that offers free routines and even a guided coach—no login needed, which I appreciate. Anyone else here use movement or gentle exercise to counterbalance screen time?

OpenAI Community 2026-07-10 20:01 UTC Score 37.0 AI-116-20260710-social-media-e716c887

Feature Request: Skills, URL Sources, and Sub-Projects for Better Long-Term Workflows in ChatGPT Projects

Thanks for taking the time to lay this out so clearly. These are thoughtful suggestions, especially for people using Projects as a long-term professional workspace. A few parts already exist in some form: Skills are available on certain plans, and Projects can use some connected sources. However, broader Plus access, arbitrary URL sources with refresh support, and nested sub-projects are still valid gaps. We’re sending this feedback to the team for logging and consideration. The examples and use cases you included should help make the request easier to evaluate. -Mark G.

Medianama AI 2026-07-10 13:21 UTC Score 36.0 USR-0211-20260710-regional-new-1a3f7d4d

Google deepens Gemini’s presence in the ads business

The 'Business Agent for Leads' is a Gemini-powered chatbot trained on an advertiser's website directly embedded inside a Search ad, allowing users to seek answers about the brand & its products and convert those into leads. The post Google deepens Gemini’s presence in the ads business appeared first on MEDIANAMA .

InfoWorld AI 2026-07-10 10:00 UTC Score 73.0 USR-0126-20260710-global-ai-ne-74596739

Meta launches low-cost Muse Spark 1.1 as enterprise AI spending comes under scrutiny

Meta has unveiled Muse Spark 1.1, saying the frontier AI model rivals leading LLMs on coding, computer use, and agentic AI benchmarks while undercutting OpenAI and Anthropic on API pricing, potentially lowering the cost of deploying AI agents in enterprises. The latest model, which was teased last week, matched or was competitive with leading models, such as Claude Opus 4.8, Gemini 3.1 Pro, and GPT 5.5, across several agentic AI, coding, and computer-use benchmarks, including SWE-bench Verified, Terminal-bench, BrowseComp, SpreadsheetBench, and OSWorld, Meta wrote in a blog post . Muse Spark 1.1, which is currently in public preview and available via the Meta Model API, will cost $1.25 per million input tokens and $4.25 per million output tokens, the company noted . By comparison, OpenAI charges $5 per million input tokens and $30 per million output tokens for GPT-5.5, while Anthropic charges $5 and $25, respectively, for Claude Opus 4.8. Google’s Gemini 3.1 Pro, on the other hand, is priced at $2 per million input tokens and $12 per million output tokens. Lower prices may open doors, not close deals That sheer difference in API pricing, according to Pareekh Jain , principal analyst at Pareekh Consulting, is enough to attract CIOs’ attention, at least for pilots, at a time when enterprises are trying to scale agentic deployments: “Pricing matters because inference costs increase rapidly when thousands of agents are working continuously.” “Output tokens are often the largest…

InfoWorld AI 2026-07-09 12:54 UTC Score 59.0 USR-0126-20260709-global-ai-ne-d6361704

JetBrains seeks to unify fragmented AI-based software development with governance suite

AI promised to make software development faster, but for many enterprises, it has also created a new management challenge: developers increasingly rely on a mix of coding assistants, AI agents, and models that operate in isolation, making it harder for engineering leaders to govern usage, share knowledge, and control costs. JetBrains has sought to address these challenges with a new suite of tools and capabilities named JetBrains AI for Teams and Organizations that supports nearly all coding tools, their respective CLIs, and most IDEs, such as Claude, Codex, Gemini, Junie, IntelliJ, Pycharm, and Rider, with support for VS Code to be added soon. The suite, which consists of capabilities like team automations and cloud agents, JetBrains Context, JetBrains Central, and JetBrains Central CLI, will allow enterprises to manage AI-assisted software development from a single control layer while allowing developers to continue using their preferred coding assistants and IDEs, Oleg Koverznev , head of agent systems at JetBrains, wrote in a blog post . While team automations and cloud agents are intended to let AI agents run long-running engineering tasks in managed cloud environments and trigger workflows based on repository events or schedules, JetBrains Context is designed to provide a shared understanding of project code, documentation, and development activity across tools, Koverznev added. Complementing those collaboration capabilities, JetBrains Central serves as the administrat…

IEEE Spectrum AI 2026-07-09 12:00 UTC Score 61.0 AI-019-20260709-global-ai-ne-7f51dbd7

Large Tabular Models Excel Where LLMs Fail

The large language models (LLMs) that form the basis of generative AI chatbots such as ChatGPT, Claude, and Gemini can generate uncannily human-like text and images. But these models still struggle with a skill that, ironically, looks at face value to be right in their wheelhouse: analyzing structured data. A new type of generative AI is set to change this situation. Although you can get your favorite chatbot to solve intractable math problems , review dense legal documents, compose a catchy pop song , or put together some slick PowerPoint slides, give it anything more than a small table and it doesn’t have a clue what to do. For most companies and organizations, the most important data sits in spreadsheets. Whether it’s a bank’s transaction logs, a marketing agency’s website metrics, clinical trial participants’ vital signs, or the vast amount of proton collision information produced at atom smashers like the Large Hadron Collider, structured, row-and-column data runs the world, and LLMs can’t deal with it. AI startup Fundamental is pioneering a new type of AI foundation model, known as a large tabular model (LTM), to fill the gap. Fundamental came out of stealth mode on 5 February 2026 with US $275 million in funding and a model called NEXUS , purpose-built for tabular data. Now, the model is being adopted by companies such as Amazon Web Services, while others race to build their own LTMs. Why LLMs struggle with spreadsheets Part of why structured data has garnered less at…

Korea AI Times 2026-07-09 07:53 UTC Score 43.0 USR-0048-20260709-global-ai-ne-d3c27ba1

구글 포토, '비디오 리믹스' 출시..."터치 몇 번으로 영화 같은 영상 변환"

전문적인 영상 편집 기술 없이 몇 번의 터치만으로 일반 영상을 영화 같은 분위기나 예술 작품 스타일로 변환할 수 있는 기능이 \'구글 포토\'에 추가됐다.구글은 8일(현지시간) 생성 AI 기반 동영상 편집 기능인 \'비디오 리믹스(Video Remix)\'를 공개했다.구글의 최신 멀티모달 모델 \'제미나이 옴니(Gemini Omni)\' 기반 기술이다. 이는 텍스트와 이미지, 영상 등 다양한 입력을 동시에 이해하고 처리하는 네이티브 멀티모달 모델로, 이번 비디오 리믹스에서는 영상 편집과 스타일 변환을 담당한다.비디오 리믹스는 구글 포토의 \'만

The Decoder 2026-07-08 14:45 UTC Score 62.0 AI-168-20260708-regional-ai--2c5a5b0b

Google Deepmind adds background execution and MCP support to Gemini API managed agents

Google Deepmind is adding four new features to Managed Agents in the Gemini API. Agents can now run asynchronously in the background, connect directly to remote MCP servers, use custom functions alongside sandbox tools, and refresh credentials without losing state. The article Google Deepmind adds background execution and MCP support to Gemini API managed agents appeared first on The Decoder .

Synced 2026-07-08 14:28 UTC Score 54.0 AI-041-20260708-ai-specialis-a8f824bb

Comment on Stanford U Demonstrates Meta-Reinforcement Agents Gain Language Skills Without Direct Language Supervision by google gemini omni

This is fascinating—it mirrors how humans pick up language through context rather than direct instruction. I’ve been following similar non-supervised approaches on Google Omni, and this Stanford work reinforces the idea that language can emerge naturally from goal-driven interaction. Curious how this scales beyond toy environments.

AI Stack Exchange 2026-07-07 17:38 UTC Score 36.0 AI-110-20260707-social-media-23b05fca

Searching for an intermediary to push AI co-discovered proof to Fermat's Last Theorem (Classical Number Theory, an easy read)

FLT Proof discovered working with Gemini Pro which uses the Alpha-Proof module written in LEAN4, and may go by the name "DeepThink" or DeepMind". [A short proof validation performed by Gemini Pro today, 7–7–2026.][1] [The short proof, coded sort of in computer programming form][2] [Spatial reasoning assistance tables][3] Note, proof link [2] is to a pdf, though I sent the AI the ODT (Libre Office file). A little easier for the AI to parse correctly. I can send the ODT file to anyone interested. [1]: https://share.gemini.google/LwBOMBEZfzBz [2]: https://fermatstheory.wordpress.com/wp-content/uploads/2026/06/docking-the-proof-viii.pdf [3]: https://fermatstheory.wordpress.com/wp-content/uploads/2026/06/p-squared-residues.pdf

InfoWorld AI 2026-07-07 01:16 UTC Score 59.0 USR-0126-20260707-global-ai-ne-f7a5d3d0

AI agents fall for indirect prompt injection traps

Some autonomous AI agents fell victim to frauds, reinforcing how easily some high-end enterprise agents can be conned by schemes that would fool few, if any, humans, Zscaler found in a test of major LLMs. The security vendor looked at various forms of indirect prompt injection (IPI) traps and found that, whereas many models fell victim to the schemes, some of the lower-level LLMs fared better than their pricier siblings. The Zscaler testing found, for example , that four models were found to be “vulnerable”: Llama3-3-70b-instruct; Llama3-2-90b-instruct; Gemini-3-flash; and Gemini-2.5-pro. Three models were found to be “safe”: Llama4-maverick; Gemini-3.1-pro; and Gemini-3.1-flash-lite. Those results indicated that the scam resistance of Gemini-2.5-pro was seemingly weaker than that of Gemini-3.1-flash-lite. But Noah Kenney , principal consultant at Digital 520, said that there is not necessarily any valuable takeaway from that revelation, because agents constantly change behavior as they feed on new data and revise their analyzed assumptions. That means an agent that failed a specific test might very well pass the identical test an hour later, he said. “The risk of an agent is constantly changing and that can cause vastly different results. You can’t assume the results are generalizable. The test result is only at one point in time,” Kenney pointed out. Zscaler “is trying to prove a point that I don’t think the data necessarily proves.” Kenney added that having a clean “safe/…

InfoWorld AI 2026-07-07 01:16 UTC Score 59.0 USR-0126-20260707-global-ai-ne-0d3954f4

Zscaler finds autonomous agents succumb to IPI traps

In a test of major LLMs, Zscaler found that some autonomous AI agents fell victim to frauds, reinforcing how easily some high-end enterprise agents can be conned by schemes that would fool few, if any, humans. The security vendor looked at various forms of indirect prompt injection (IPI) traps and found that, whereas many models fell victim to the schemes, some of the lower-level LLMs fared better than their pricier siblings. The Zscaler testing found, for example , that four models were found to be “vulnerable”: Llama3-3-70b-instruct; Llama3-2-90b-instruct; Gemini-3-flash; and Gemini-2.5-pro. Three models were found to be “safe”: Llama4-maverick; Gemini-3.1-pro; and Gemini-3.1-flash-lite. Those results indicated that the scam resistance of Gemini-2.5-pro was seemingly weaker than that of Gemini-3.1-flash-lite. But Noah Kenney , principal consultant at Digital 520, said that there is not necessarily any valuable takeaway from that revelation, because agents constantly change behavior as they feed on new data and revise their analyzed assumptions. That means an agent that failed a specific test might very well pass the identical test an hour later, he said. “The risk of an agent is constantly changing and that can cause vastly different results. You can’t assume the results are generalizable. The test result is only at one point in time,” Kenney pointed out. Zscaler “is trying to prove a point that I don’t think the data necessarily proves.” Kenney added that having a clean “…

Roboflow Blog 2026-07-07 00:07 UTC Score 33.0 USR-0088-20260707-ai-specialis-c6ae9e63

Predictive Maintenance with Vision AI

Use RF-DETR to detect early bearing defects and Gemini 2.5 Pro to analyze defect severity, explain potential failure causes, and generate maintenance recommendations.