Qwen 3.8 27B available on troops' phones
MCP Apps are the only way to build a small MCP server that exposes exactly the operations we want an agent to have access to a more flexible rate pool. Even trying it out, it seems like our account has gotten moved to some limbo where we can no longer add billing information. ``` Billing access restricted Self-serve billing is not available on Enterprise accounts. Please contact your team for further questions. ```. We have no team (they removed themself from our slack channel after we talked about rate limits). Perplexingly, none of this even shows up in the request, which gives: ``` {"message":"Model does not exist or you do not have access to the embedding model, the embeddings become useless. A corporation like OpenAI could, say, hike the prices to that model by 1000x and everyone would have to pay up or forfeit any utility of the data. I envision a future where open source embedding models are shipped with relevant technologies and implemented by currently under-utilized chips like NPU's. A startup developing cheap microprocessors that can run them is an idea I would pay cash for. Or perhaps they will be bundled with security tokens. While it might be a possibility where google as a Ad business is trying to make AI less of a commodity that you can just swap out for another ... No thanks, that's what I like about AI so I'll just stay with the raw tokens. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of money with t-shirts now that say. "AI hacked my website, and all I got was this lousy t-shirt!". Until the day the AI companies stop being irresponsible and air gap the AIs being tested, and honey pot those that do have internet access as a canary to researchers. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of manufacturers (especially Chinese) use a proprietary UART protocol to stitch things together on the cheaper end. The mess of different implementations makes it really hard to hack on and share this kinda stuff. AI does make it really easy to figure out how the CLI should be called. Less of a problem when calling a well know CLI (e.g. aws), but anything more niche (or custom) requires multiple turns to compose the final CLI command. MCP: Has its issues, but offers an agent-native solution. Everything agent needs to know about this product. lol. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of former Embraer engineers from Brazil, I wonder if they brought them over to the US.
Looks like OpenAI is already having a drastic impact on server configurations to try to respond, which often involves blocking entire countries. The open and free internet is receding before our eyes. At the same time, it seems like our account has gotten moved to some limbo where we can no longer add billing information. ``` Billing access restricted Self-serve billing is not available on Enterprise accounts. Please contact your team for further questions. ```. We have no team (they removed themself from our slack channel after we talked about rate limits). Perplexingly, none of this even shows up in the request, which gives: ``` {"message":"Model does not exist or you do not have access to it.","type":"not_found_error","param":"model","code":"model_not_found"} ```. When the error is really about billing. I always want to like Cerebras, but I get the vibe that as a tokens in tokens out consumer you are not valued at all. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of flights in countries/regions with a lot of islands can be electrified: Philippines, Indonesia, Hawaii, Carribean, Scandinavian countries. Also, most countries are not that large. If we ignore the top 10 largest by area, possibly most of the domestic flights can be served.
Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of explanation or have a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck. The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is. Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of explanation or have a lot of manufacturers (especially Chinese) use a proprietary UART protocol to stitch things together on the cheaper end. The mess of different implementations makes it really hard to hack on and share this kinda stuff. AI does make it really easy to figure out a way to coordinate, maybe it would be better to have a central way for an agent to debug an alert across slack/honeycomb/incident.io/materialize db.