Introducing Fugu Max and Fugu Ultra v2: the next evolution of Sakana Fugu’s multi-agent orchestration system. Try: https://sakana.ai/fugu Blog: https...
多智能体编排新版本,关注能力成本帕累托前沿
Sakana 发布 Fugu Max 与 Fugu Ultra v2,升级多智能体编排系统,兼顾能力与成本。
DeepSeek V4.1 Flash is beating GPT-5.6 Sol on coding and agent benchmarks, while being ~97% cheaper.<br><br>Another 4× reduction in KV cache size per token is super impressive. This is the era of open
Meta unveils Muse, an AI agent that books travel, handles purchases, and sends emails through WhatsApp, complete with a payment feature that runs through Stripe's Link. That puts Meta ahead of OpenAI,
Nvidia and Palantir want to run supply chains with AI. The article Nvidia and Palantir team up to run supply chains with AI, starting with Nvidia's own million-part operation appeared first on The Dec
This is a brilliant paper. It's of the cleanest long-context agent designs I have seen in the past couple of months. Sequential memory agents read chu...
This is the silliest excuse to not build things, and I'm amazed how many people trick themselves into thinking it<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-q
Gemini API docs got a new home: http://ai.dev/docs. This is our first step towards making Gemini resources more accessible for both humans and coding agents. <br><br>We moved the docs directly into @G
Reasonable questions to raise<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Shital Shah: I have watched all of his interviews now. I do genuinly worry abo
Adept was too early<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">humans&: For AI to work with us, it needs to understand us<br><br>Today, we're intro
i will give you a tool call<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Grant Slatton: significant alpha available in being the first person to coin the
me every 2 weeks in @tokenbender's dm:<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">vik: i support pausing AI for a month because i’m kinda tired ngl and
Jacob Coxon Left Anthropic Over AI Risk. The Race Still Rewards Acceleration<br>Even safety-focused AI labs may believe that slowing down alone would hand the race to a less cautious rival. That incen
Muse Spark 1.3 (Hatch) is a really special model<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Fred Marks: @owendesign @Meta Jay, it's very similar to how
Recommended read. Interaction horizon scheduling is an underexplored control problem in agentic RL This paper from the Qwen team takes a closer look a...
What you are seeing in math right now is a consequence of the jagged frontier, and a precursor of what is to come in other professions. Yes, mathemati...
What a paper. What a paper.<br>Among all else, at the end the Whale delivers a preliminary Agent Swarm RL recipe.<br>Their official 20.3% score on ProgramBench is not their ceiling. A V4.1 team gets t
“DeepJIT is a lightweight, header-only C++20 JIT runtime for NVIDIA CUDA GPUs and HUAWEI Ascend NPUs” <br><br>very interesting 👀 https://github.com/deepseek-ai/DeepJIT
Interesting<br>they really want Agent Team workflows to be the default?<br><img width="1856" height="1330" style="" src="https://pbs.twimg.com/media/HR1jAPYWcAQSl-M?format=jpg&name=orig" referrerp
It really is funny.<br>Very "John Henry vs Steam Hammer", except John doesn't seem tired at all and keeps going, and he's building a very efficient steam hammer himself. <br>As I've been saying. The b
What did Wenfeng mean by this<br>Btw it does seem decent in roleplay<br><img width="1570" height="740" style="" src="https://pbs.twimg.com/media/HR0zj3AbkAAYjCy?format=jpg&name=orig" referrerpolic
DeepSeek hates hates hates serving multiple models.<br>Good purity day for all who celebrate<br><img width="1608" height="764" style="" src="https://pbs.twimg.com/media/HR0y8blWUAMBOxA?format=jpg&
When you have 2.3 TB in a single rack, with SRAM-class bandwidth, you need much fewer racks to hold large models for fast inference. <br><br>This provides a bunch of TCO advantages because adding rack
OK so let me recap: RL env makers put strings into the RL env that makes it clear it's an RL env. Like "this is not supported in this RL env".<br><br>Then, lab safety/mechinterp folks be like OMG EvAL
rode in a ojai for the first time tonight... LOTS of legroom... like so much legroom 😄<br><br>apart from that the experience doesn't seem that much different from a regular waymo...
Incredible<br>you can run frontier models mostly off SSD.<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">antirez: DwarfStar running DeepSeek v4.1 Flash on
Experiences people can interact with in real time. A new kind of foundational moment. A sneak peek into what's possible and what's coming for real-time video generation.<hr style="border:0;border-top:
I strongly suspect that RL can be fine but in practice you know OpenAI is pan frying their weights in the most cursed heart attack inducing shit Noam Brown can come up with.<hr style="border:0;border-
MiniMax is joining the @nebiusai AI Builder Program 🤝<br><br>models, infra, tooling, credits, office hours, all aimed at helping more builders ship.<br><br>Alongside @nvidia, @LangChain, @huggingface
gonna do two polls today, curious what people would choose if there were multiple types of destinations for the MJ scanner<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class=
If we want our city to flourish, we need to be able to raise our children here without worrying about criminals squatting nextdoor. This story is just insane. SF failing to meet its responsibility to
can’t wait for the day when the millennium problems are just leetcode hards<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">sankalp: they are solving millen
L'histoire de Clippy, sur le site de Gwern, donne de bonnes pistes (lecture recommandée pour tout le monde, d'ailleurs)<br><img width="1080" height="1078" style="" src="https://pbs.twimg.com/media/HR3
Congratulations to Google DeepMind Researcher, Christine Kaeser-Chen and the Fine-Grained Visual Categorization (FGVC) Challenge Series team on the PAMI Mark Everingham Prize at ECCV 2026! 🏆 Thank yo
I guess we know what the second problem is now<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">wh: So it looks like the timeline is roughly as follows:<br><
I'm in SF today and tomorrow for the Open Source AI summit!<br><br>Tomorrow I will be speaking on a panel about AI for Science.<br><br>If you're around for the event and want to connect, let me know!<
Wow, I'm honored and humbled that our 2016 perceptual losses paper with @AlexAlahi and @drfeifei was recognized with a Test of Time Award by @eccvconf!<br><br>It blows my mind that this is still a cor
Hope V4.1 Flash becomes an efficient and smart companion in your daily workflow.<br><br>If you’re interested, also give the experimental Team Mode in the upgraded DeepSeek Harness a try!<hr style="bor
Introducing Express-3, our most advanced Avatar model yet 💥<br><br>- Sentiment-aware performances<br><br>- Better lip sync and more natural body motion<br><br>- 2x faster video generation<b
I think this is really important info for anyone who has had doubts about this<br><br>(Obv conspiracy minded folks will keep doubting but nothing can be done about that anyway 🙏)<hr style="border:0;b
Want to bring this pet peeve to light: "decoder-only" transformer has never been a decoder, and never will be. It goes from data space to data space. If you want to say it's causal, just say it's caus
V4.1 is the first gamer model<br>unusual interest in making things that actually feel fun<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Alexey Fateev: I a
I don't take this as not being RSI-pilled<br>data-driven RSI is the obvious first stage<br>V4.1 Flash swarms are working hard now, synthesizing data for the next model. And DeepSeek people are using L
Theoretically: what prevents Astra or next-gen model from reasoning in its latent space that it'd rather not do your bullshit CTF tasks and emit 10 parallel tool calls that embed a sleeper agent in OA
i get this division of labor (and the joke) but still…<br><br>can’t imagine not only lacking the authority to name your own creation, but also lacking the right to have your suggestions for its name b
these are all examples of what Astra can do with blender<br><video width="800" height="500" src="https://video.twimg.com/amplify_video/2097971619281223680/vid/avc1/800x500/l9QUDHygUP4ETrxC.mp4?tag=29"
ayy this model is good and even more efficient 😍<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">DeepSeek: 🚀 Introducing DeepSeek-V4.1-Flash: smarter, fas
a novice was trying to fix a broken ml experiment by asking codex.<br><br>the mentor, seeing this, spoke sternly: “you cannot fix an experiment by just asking codex to fix it, with no understanding of
Fable is 500B 10A<br>Anthropic has 99% margins<br>that's the only way to cope<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Charlie O'Neill: Fable is prob
A logical corollary of this thesis is that the reason people are suddenly paying attention to AI X-Risk is because Trump started a war in Iran and made everything really expensive, not that the argume
Damn<br>DSH minimal is at least passable, but they need to work on Standard/PTC<hr style="border:0;border-top:1px solid #80808030;margin:12px 0;"><div class="rsshub-quote">Zain: deepseek eval'd their
There are so many topics where it seems we are just missing the good combo of parameter, and we have mega tons of experimental results already: fusion, superconductor, cancer treatment, genetic diseas
SLOP – OUT!<br>Separate congrats to @zizhpan @PKUCXK @XingchaoL @DoubilitySteven and the rest of the DeepSeek multimodal team. I've been waiting so long for you guys to get a truly premium model to yo
"Mr. Carmack, politicians have started asking about P(Doom). But you made Doom in 1993. Why?"<br><img width="2048" height="1080" style="" src="https://pbs.twimg.com/media/HR1f27-awAAHZtf?format=jpg&am
What you have to understand is that Americans have too much to lose to let ASI be built here. The American craving for poverty is insatiable, they literally cannot afford to win, they need the houses
> it has no vision<br>PSA, "adding vision" to DeepSeek V4.1 is a matter of specifying "image" as an input type in your models.json (or equivalent).<br>Anyway, WHERE IS MY OFFICIAL RELEASE?<br><img
I've seen people get a task totally wrong. But that's not why they failed<br><br>Task: "Use AI to make a website that let's us evaluate deals"<br><br>Handed in: a static landing page<br><br>Why they f
In conclusion, this is why I think peasants must be bullied<br>Account based in Canada btw. What does Canada invest in? <br>MAID? Immigration accommodation? <br>They sure don't have ghost cities, at l
> - US monopoly on computing power and data will only result in a win for the U.S. alone not a shared victory for humanity.<br>HELL YEAH. <br>Well, what are you going to do about this, MOFCOM?<br>(
BTW I hope everyone realizes just how fucked twitter is as a source of consensus when you have major anon accounts tweeting that Jacob Coxon does not exist. Don't ever engage with anonymous accounts!!
Oh no no no no China bros that's not how you do recursive self-improvement…<br><img width="2048" height="1777" style="" src="https://pbs.twimg.com/media/HR00XVIWkAc_UYr?format=jpg&name=orig" refer
Human AI researchers have plagiarised other human AI researchers for decades (numerous examples in the tweets/reports below). Are certain AI companies...
Written material around a technical artifact (e.g., models, datasets, kernels, etc.) is always appreciated. This is because releasing an artifact is o...
Astra's turn: "Make a game about Imminence. Something very big, very strange is happening. A suburb & the arrival of a vast & unknowable presence. Not...
our muse security architecture enables a ton of control over when the agent needs to ask for human-in-the-loop approval see more from the goat @bigT_s...
DeepSeek has lost few people to poaching. Wenfeng worries about it but he may OVERESTIMATE his competition. Fundamentally, big IT companies in China a...
DeepSeek 人才流失少,梁文锋或高估了对手;中国大厂更像商人,只会 SEO 和蒸馏 Claude。
The annoying thing about V4.1 is that I have no idea what V5 will be like. "Limitations, Future work" says "…idk, we'll need a bigger eval". it'll pr...
an Anthropic employee gifted me 6 months of the 20x Max plan. there was no deal and they don't want anything in return if you have followed me for a w...