Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🐝
Article: Where the Hivemind Comes From
Burton Lancaster
PRO
RiverRider
6
4
25
Follow
LiteMind's profile picture
James56968's profile picture
SuperPauly's profile picture
40 followers
·
46 following
https://sunstonenorth.com
Space-Bacon
AI & ML interests
Explainable AI
Recent Activity
liked
a dataset
about 12 hours ago
detection-datasets/coco
posted
an
update
about 13 hours ago
SWE-bench Verified scores whether an agent's patch passes the tests. It does not score whether the agent found the right file first, which is the step before it. We measured that step on all 500 instances. A 33M-parameter encoder, BAAI/bge-small-en-v1.5 at 384 dimensions, names the correct file first for 229 of 500 (0.458). Plain text search over the same checkouts gets 35 (0.070). The number that makes those readable is the floor. Hand the same index a bug report from an unrelated project and it still lands the gold file at rank 1 for 5 of 500 (0.010). So 0.458 is 45.8x chance, not 45.8x nothing. That ratio is where the argument is. recall@50 reads 0.954 and sounds like a solved problem. An unrelated report reaches the same top 50 for 0.244 of instances, so the margin over chance falls from 45.8x at k=1 to 13.6x at k=10 and 3.9x at k=50. The headline that looks best is the one carrying the least. All 500 ranked lists are published under CC BY 4.0, so the floor can be recomputed rather than believed. The article also ends with eleven corrections to claims we made earlier and got wrong, including one where the lever we proposed turned out to cost accuracy rather than buy it. We have not measured the patch step. This is the one before it. For anyone who doesn't live in SWE-bench: it gives model a real GitHub issue from a real project and scores whether the code it writes makes that project's tests pass. That single score covers two jobs, finding the file that needs changing and then changing it correctly, and only the pair is ever scored. The first step is what we measured. Nothing here writes code or runs a test, so 0.458 is a hit rate for naming the right file, not a SWE-bench resolve rate. Article: https://huggingface.co/blog/RiverRider/finding-the-file-localisation-on-swe-bench-verifie Data: https://huggingface.co/datasets/RiverRider/swebench-localisation
replied
to
their
post
1 day ago
Black Window — a chat model in your browser tab, on your hardware. A memory that stays on the device that opened the page. https://blackwindow.xyz Open the site, pick a model (about 0.6B to 8B), hit Load. The weights run in that tab, on that computer. After they load, the network can drop. The context window is a working set, auto-sized to that device, up to ~32K tokens. Behind the window is the Weave. Every file, picture, recording, link, lookup, and reply is embedded as it arrives. Drop in audio and it is transcribed. Drop in an image and it is described. A question pulls the nearest passages back as notes. A long document is walked once so later questions can use the whole file, not the first pages. Nothing leaves that tab unless you turn on live lookup or connect a rented GPU box, and the chat says so each time. Prompts can go to the box. Files and the Weave stay in the tab. Console on that page: bw.ask, bw.search, bw.digest, bw.notes. A local relay exposes /v1/chat/completions on localhost so other tools on the same computer can talk to the tab. The tab polls the relay. That is the boundary. Not a server with a policy. Your hardware, a window, a Load button. If on mobile add to home-screen for best performance. If you break it lmk. It can serve a few hundred of you at a time before I have to buy a real server.
View all activity
Organizations
RiverRider
's datasets
14
Sort: Recently updated
RiverRider/swebench-localisation
Viewer
•
Updated
1 day ago
•
500
•
590
RiverRider/srt-hivemind
Updated
11 days ago
•
654
•
2
RiverRider/srt-omni-crossvendor-states
Updated
16 days ago
•
197
RiverRider/srt-cxr14-frozen-probe
Updated
18 days ago
•
495
•
2
RiverRider/srt-nla-gemma4-artifacts
Updated
19 days ago
•
133
RiverRider/srt-depth-probe-artifacts
Updated
19 days ago
•
155
RiverRider/srt-omni-manifest
Viewer
•
Updated
20 days ago
•
1
•
71
RiverRider/srt-qwen38-coco-states
Updated
24 days ago
•
93
RiverRider/srt-coco-thumbs
Viewer
•
Updated
24 days ago
•
10.3k
•
1.62k
RiverRider/srt-nla-gptoss20b-artifacts
Updated
Jul 2
•
178
RiverRider/srt-nla-targets-gemma2-2b-v1
Updated
May 21
•
24
RiverRider/srt-nla-targets-llama32-3b-v1
Updated
May 21
•
20
RiverRider/srt-nla-targets-v1
Updated
May 18
•
55
RiverRider/zoolander-corpus-v23
Viewer
•
Updated
May 5
•
93.3k
•
21