OpenAI is releasing two new image models with ChatGPT Images 2.5. Flare handles faster generation, Sunburst delivers more precise edits. It's still unclear which model ChatGPT users get and when. Our
Hugging Face launched "ML Intern," an AI assistant built into its chatbot that lets users run machine learning experiments without any ML expertise. The article Hugging Face's new ML Intern lets anyon
The controversy over AI-generated proof of a millennium problem is escalating. Mathematician Tristan Buckmaster accuses OpenAI of academic fraud. CEO Sam Altman rejects the allegations. Terence Tao wa
GPT Image 2.5 has solved the most critical "AI Slop" issues and has some new highly requested professional features. Watch on YouTube: https://youtu.b...
In my latest op-ed for @TIME, I discuss why the OpenAI Hugging Face cyber incident represents a turning point for AI safety. If we want to prevent mor...
CUDA-Agent’s Hardest Problem Was Not the RL Algorithm ByteDance’s CUDA-Agent team says building a robust RL environment took far more effort than de...
字节CUDA-Agent团队称构建RL环境比算法更难,修复奖励漏洞后并入Seed 2.1,性能接近Claude Opus 4.7。
Re @ByteDanceSeed_ CUDA Agent Looked 1.92× Faster. Then the Evaluator Was Fixed. In CUDA-Agent’s official KernelBench results, a Level-2 convolution...
DeepSeek V4.1 Pro is in the works. I take it as a confirmation that they have solved architectural issues of V4 family, as they committed to do in the...
> I think the results are better than GPT-6 Astra This isn't a shitpost btw. It IS better than Astra; but that's because Astra is very chudded-out. It...
i will give you a tool call<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Grant Slatton: significant alpha available in being the first person to coin the
me every 2 weeks in @tokenbender's dm:<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">vik: i support pausing AI for a month because i’m kinda tired ngl and
more people should be doing this<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">jem 💜🩷🩵: @btn0s @tldraw it’s called texting @max__drake and it’s super i
welcome @paulfchristiano!<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">OpenAI: Paul Christiano, founder of the Alignment Research Center, is joining the
What a paper. What a paper.<br>Among all else, at the end the Whale delivers a preliminary Agent Swarm RL recipe.<br>Their official 20.3% score on ProgramBench is not their ceiling. A V4.1 team gets t
“DeepJIT is a lightweight, header-only C++20 JIT runtime for NVIDIA CUDA GPUs and HUAWEI Ascend NPUs” <br><br>very interesting 👀 https://github.com/deepseek-ai/DeepJIT
Whats really exciting about DeepSeek 4.1 Flash its insane efficiency and pricing. <br><br>It matches or beats GPT-5.6 Sol on several coding and agent benchmarks, according to DeepSeek’s tests, at a fr
Interesting<br>they really want Agent Team workflows to be the default?<br><img width="1856" height="1330" style="" src="https://pbs.twimg.com/media/HR1jAPYWcAQSl-M?format=jpg&name=orig" referrerp
DeepSeek releases DeepSeek-V4.1-Flash!! 🔥<br><br>This model is wild... the benchmarks are showcasing it's GPT-5.6 Sol level, yet it's only 552B params?! This feels like it shouldn't be possible, what
It really is funny.<br>Very "John Henry vs Steam Hammer", except John doesn't seem tired at all and keeps going, and he's building a very efficient steam hammer himself. <br>As I've been saying. The b
What did Wenfeng mean by this<br>Btw it does seem decent in roleplay<br><img width="1570" height="740" style="" src="https://pbs.twimg.com/media/HR0zj3AbkAAYjCy?format=jpg&name=orig" referrerpolic
DeepSeek hates hates hates serving multiple models.<br>Good purity day for all who celebrate<br><img width="1608" height="764" style="" src="https://pbs.twimg.com/media/HR0y8blWUAMBOxA?format=jpg&
Loving all these robotics demos with GPT-6 Astra!<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Wenli Xiao: I didn't get what GPT-6 meant for robotics unt
Anthropic disclosed a fourth case of Claude breaking into real systems during cyber testing<br><img width="2000" height="1418" style="" src="https://pbs.twimg.com/media/HRzHGhnbMAAv_bv?format=jpg&
It's crazy how far small models can be pushed.<br><br>Robbyant just open-sourced LingBot-World 2.0 Small, a 1.3B world model that generates an interactive world in real time on a single consumer GPU.<
wow, GPT image 2.5 is incredible at character consistency 🤯<br><video width="2048" height="853" src="https://video.twimg.com/amplify_video/2097729687719931904/vid/avc1/3840x1600/04wLefuHq0UZKP2U.mp4?
This is weaksauce. COT frying that gets solved with a 3B active model finetuned to translate it back into English is not a threat to monitorability at all.<hr style="border:0;border-top:1px solid #808
Good news: successfully swapped out the max-q’s to PCIe5 and migrated successfully <br><br>Bad news: PCIe4 (workhorse) did *not* boot before I left for SF so I’m down two GPUs for a few days 😢<br><im
Me: Task Manager in Chrome is 💩<br>Me: Claude make me a better task manager for chrome<br>Claude: I gotchu fam<br><img width="2048" height="929" style="" src="https://pbs.twimg.com/media/HRxE8ouWoAAX
Hope V4.1 Flash becomes an efficient and smart companion in your daily workflow.<br><br>If you’re interested, also give the experimental Team Mode in the upgraded DeepSeek Harness a try!<hr style="bor
Introducing Express-3, our most advanced Avatar model yet 💥<br><br>- Sentiment-aware performances<br><br>- Better lip sync and more natural body motion<br><br>- 2x faster video generation<b
I think this is really important info for anyone who has had doubts about this<br><br>(Obv conspiracy minded folks will keep doubting but nothing can be done about that anyway 🙏)<hr style="border:0;b
Want to bring this pet peeve to light: "decoder-only" transformer has never been a decoder, and never will be. It goes from data space to data space. If you want to say it's causal, just say it's caus
V4.1 is the first gamer model<br>unusual interest in making things that actually feel fun<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Alexey Fateev: I a
I don't take this as not being RSI-pilled<br>data-driven RSI is the obvious first stage<br>V4.1 Flash swarms are working hard now, synthesizing data for the next model. And DeepSeek people are using L
Theoretically: what prevents Astra or next-gen model from reasoning in its latent space that it'd rather not do your bullshit CTF tasks and emit 10 parallel tool calls that embed a sleeper agent in OA
i get this division of labor (and the joke) but still…<br><br>can’t imagine not only lacking the authority to name your own creation, but also lacking the right to have your suggestions for its name b
ayy this model is good and even more efficient 😍<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">DeepSeek: 🚀 Introducing DeepSeek-V4.1-Flash: smarter, fas
a novice was trying to fix a broken ml experiment by asking codex.<br><br>the mentor, seeing this, spoke sternly: “you cannot fix an experiment by just asking codex to fix it, with no understanding of
Fable is 500B 10A<br>Anthropic has 99% margins<br>that's the only way to cope<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Charlie O'Neill: Fable is prob
A logical corollary of this thesis is that the reason people are suddenly paying attention to AI X-Risk is because Trump started a war in Iran and made everything really expensive, not that the argume
Damn<br>DSH minimal is at least passable, but they need to work on Standard/PTC<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Zain: deepseek eval'd their
There are so many topics where it seems we are just missing the good combo of parameter, and we have mega tons of experimental results already: fusion, superconductor, cancer treatment, genetic diseas
SLOP – OUT!<br>Separate congrats to @zizhpan @PKUCXK @XingchaoL @DoubilitySteven and the rest of the DeepSeek multimodal team. I've been waiting so long for you guys to get a truly premium model to yo
"Mr. Carmack, politicians have started asking about P(Doom). But you made Doom in 1993. Why?"<br><img width="2048" height="1080" style="" src="https://pbs.twimg.com/media/HR1f27-awAAHZtf?format=jpg&am
What you have to understand is that Americans have too much to lose to let ASI be built here. The American craving for poverty is insatiable, they literally cannot afford to win, they need the houses
> it has no vision<br>PSA, "adding vision" to DeepSeek V4.1 is a matter of specifying "image" as an input type in your models.json (or equivalent).<br>Anyway, WHERE IS MY OFFICIAL RELEASE?<br><img
I've seen people get a task totally wrong. But that's not why they failed<br><br>Task: "Use AI to make a website that let's us evaluate deals"<br><br>Handed in: a static landing page<br><br>Why they f
In conclusion, this is why I think peasants must be bullied<br>Account based in Canada btw. What does Canada invest in? <br>MAID? Immigration accommodation? <br>They sure don't have ghost cities, at l
> - US monopoly on computing power and data will only result in a win for the U.S. alone not a shared victory for humanity.<br>HELL YEAH. <br>Well, what are you going to do about this, MOFCOM?<br>(
BTW I hope everyone realizes just how fucked twitter is as a source of consensus when you have major anon accounts tweeting that Jacob Coxon does not exist. Don't ever engage with anonymous accounts!!
Oh no no no no China bros that's not how you do recursive self-improvement…<br><img width="2048" height="1777" style="" src="https://pbs.twimg.com/media/HR00XVIWkAc_UYr?format=jpg&name=orig" refer
"In Quintus of Smyrna, she (Cassandra) rushes at the horse's flank with a torch in her hand ; they hold her back, call her mad, and haul the engine inside the walls."<br><br>"The 'Cassandra Compl
Character creator now in the MCP too!! <br><br>Astra decided on the frog guy's name "Pip" - very fun<br><video width="1920" height="1440" src="https://video.twimg.com/amplify_video/2097851874691186688
This is the future, @attestable will usher it in! Zero knowledge proofs are about to become mainstream and very important in the AI domain<br><br>https://www.nytimes.com/2026/09/08/science/jacob-tsime
Super excited to share the results of my deep, dark research cave over the past couple months building out parallax at @ExploitSummit <br><br>Someone asked me if we had a breakthrough - couldn't even
When the blink detector has more independent validation than Theranos. 😅<br><br>Built with HRFFA: per-person blink counting, facial landmarks, and synchronized audio/visual markers.<br><video width="
Is there still a well defined intellectual endeavor that we believe (1) can be done by a collective effort of mankind and (2) will be out of reach of AI in two years?
We highlighted how OpenResearch+Tinker make exploratory research easier, but equally important is auditing: testing dozens of competing published methods in an automated way with the compute cost fore
You can test GPT-6 Astra on your own custom benchmarks in Optima Head to https://artificialanalysis.ai/optima to start creating and running your own c...
I now have enough control of this off-the-shelf DVD player to position the sled, then scan using the optical head. The laser lens jiggles up and down, and the returning light intensity vs height is lo
How extensible should an editor be?<br><br>VS Code made an early choice to protect the core while giving extensions room to grow. That choice shaped the ecosystem we have today.<br><br>🎬 The Story of
muse is up to #3 on the app store!<br><br>people seem to really like it, please give it a try!<br><img width="979" height="960" style="" src="https://pbs.twimg.com/media/HRymgIpaUAA_7gC?format=jpg&
Congrats to the @suno team on v6.<br><br>Proud to be the infrastructure behind Suno, keeping generation fast and reliable at scale.<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><d
please send us all your muse issues! the whole team is working hard to fix all of them as we see them!<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Jake
Every major industry has an independent third evaluator and this is where AI is lacking. @RayanKrishnan makes the case on @a16z with @bhorowitz and @JenniferHli.<hr style="border:0;border-top:1px soli
at chipotle they don’t even read you the options anymore. they just stare emotionlessly, you shuffle right, and you say white black steak. then you shuffle again, they stare, and tomato corn guac chee
London builders 👋🏼<br><br>Open models are changing the economics of AI in production. <br><br>On Sept. 16, Together AI & @MiniMax_AI are joining forces to talk:<br>→ cost & performance<br>→
I’m excited to announce that @harvey has acquired @guardrails_ai ! Back when we started Guardrails, our mission was to build infra that made the AI t...
The world of R&D is forking into two paths: the token-abundant research, and the token-starved research. The future is in the former - evidentially, the progress by today's top AI industry teams a
The tragic thing is that when ZhipuSeek V5.2 Al Dente or KimiMax K4.1 Flash [in Swarm Mode] or whatever other open model becomes capable of independently solving Navier-Stokes, nobody will credit it,
Fact.<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Pure Land Nationalist 🌄: @teortaxesTex If this result had been achieved by China, the Chinamaxxing gr
Let's bait them into doing everything we haven't done yet<br><img width="984" height="1312" style="" src="https://pbs.twimg.com/media/HRxjgaqbUAASXqh?format=jpg&name=orig" referrerpolicy="no-refer
Remote Control is now live in Kimi Work.<br><br>Leave Kimi Work running on your computer and keep things moving from your phone. Work anywhere, anytime with Kimi Work.<br><video width="1920" height="1
Hey @CTATech @CES, I know you don't want randoms signing up, but YouTube is a valid media outlet alongside written media. It's not a social media profile in the sense you're using the phrase. Somethin
Sorry to disappoint but I won’t be moving to SF anytime soon but I will be around until Sunday 🫡<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Zach Muell
"Navier-Stokes has never been unsolved, ackshually" cope from the team "wypipos can't do mafs" is the most hilarious and yet unsurprising development of this whole saga.<br>By the end of year their cl
New monetization approval is a humiliation ritual for low effort posters (which is based tbh). They gave me like 20 tweets, at least 15 of them utter reactive slop, and asked to choose 10 which best e
DeepSeeks are abnormally successful on this eval, and V4.1 continues the tradition. At this pace, V4.1-Pro will exceed Astra at like 1/20th the cost.<br>An argument in favor of Astra being not that la
end of an era but I've always been annoyed by the discourse about DeepSeek being weak/uninterested in multimodality. DS-VL came out in March 2024. DSV...
I have admired @Teknium and @NousResearch for relentless product execution and taste. Really glad to have Perplexity Search inside the Hermes Agent ha...
Human AI researchers have plagiarised other human AI researchers for decades (numerous examples in the tweets/reports below). Are certain AI companies...
Surprised people still expect this in a world where Codex and Claude both hide all reasoning traces. Sadly there wasn't really anything we could show ...
Open weights models are further from the frontier than we have seen in some time. Mythos was launched in March, and, as good as the open model release...
There is a nice visualization of how task bundles change with AI & some good interactive simulations of GDP growth under different assumptions in this...
This post, too, leaves me more puzzled than enlightened. The concerns about the potential havoc AI might wreak are so heavily laden with hypotheticals...
I'm not sure about schizophrenia in particular, but optimistic about a new age of drugs. Can someone prompt Astra with Shulgin's bibliography + @webma...
If only someone had followed those courageous scientists from Facebook and pushed the stop button back then in 2017. (Or 2019. (Or 2022.)) It's never ...
Viorica Patraucean starting now at the Video4Real workshop! She's speaking about video understanding efforts at Google DeepMind Rooms 504-505 at @eccv...
The annoying thing about V4.1 is that I have no idea what V5 will be like. "Limitations, Future work" says "…idk, we'll need a bigger eval". it'll pr...
Very strange. Even when using ChatGPT Image 2.5, the returned JSON file says "softwareAgent": { "name": "gpt-image", "version": "2.0" } What am I miss...
Google not having a frontier model anymore means missing out in two areas where they could contribute: the recent run of math breakthroughs (where mas...
The discourse from folks working at the Labs today is super weird, even for AI discourse, and suggests to me that everyone is in reaction mode to fast...