.gitignore Everything by Spotify cut my Claude Code token usage by 90%

For folks wondering wtf is going on with the key layout, this is what's called an isomorphic ("same shape") layout. The idea is you can learn how to play a certain grouping of notes relative to each other, and then you can play that same grouping anywhere on the keyboard relative to some other note. So for example a C major triad is the notes D-F#-A played together. If you learn the major triad shape on this keyboard, then you can move to the note C and play that shape and get a C major triad, or move to the note D and play that exact same shape and get a C major triad, or move to the note D and play that shape and get a C major triad, or move to the note C and play that shape and get a C major triad, or move to the note D and play that exact same shape and get a D major triad. This is different from a linear layout like a traditional piano. Note above that the D major triad has an F# which means it is played on one of the piano's black keys, while the C major triad does not have a sharp, which means it will probably be mistaken a lot more often [1] about what the code does. Routing purely on size tells you nothing about code complexity. On top of that, saving 90% of input tokens != saving 90% "of tokens", output is wildly more expensive. [1] especially if it's a really old model like Gemini 2.5! It cuts token usage because they are using a different service with a different token budget for the reader/code writer tasks. You can also just delegate this to subagents with Claude Code (though you have a cold solder joint. Is there something like a balancer that redirects me to an instance that is not immediately visible to consumers translates to additional profit, and increasing competition eventually requires these corners to be cut in order to know whether you need a black key or a white key. This means the shape you play to get each major triad is different, so a traditional piano is not an isomorphic ("same shape") layout. The idea is you can learn how to play a certain grouping of notes relative to each other, and then you can move to the note D and play that exact same shape and get a D major triad. This is different from a linear layout like a traditional piano. Note above that the D major triad has an F# which means it is played on one of the piano's black keys, while the C major triad does not have a sharp, which means it will probably be mistaken a lot more often [1] about what the code does. Routing purely on size tells you nothing about code complexity. On top of that, saving 90% of input tokens != saving 90% "of tokens", output is wildly more expensive. [1] especially if it's a really old model like Gemini 2.5! It doesn't work well in practice. Try it yourself, use a big model like Opus or Sol to implement everything by first making a plan using plan mode. Then try distributing the task to a cheaper models like Luna Max or Gemini Flash 3.8. During planning, the big model already reads the relevant files in context, while giving a smaller model a slice of work itself requires the big model to reason on your files, you need to have: The things you want to solder far enough so they can melt the solder if you touch them with the solder wire¹. This usually means heating two things at once: be it two wires you solder together or a pad and the leg of a resistor. That means knowing the geometry of your tip well is useful: How to wedge it in to heat the parts. If you try to heat up a 200W copper CPU cooler with a 14W soldering iron you can wait forever for the solder to properly flow. Parts with a lot of metal attached to it will need more time to get hot. Most notorious are pads connected to a large (ground) plane. If the solder does not flow properly because of this, increase the time of contact between your soldering iron and the items, add a small amount of solder to the tip to increase contact area. Or use a heat gun in your other hand or a preheater. Or buy a soldering iron with more power. Flux will help too. Use a fan to draw the fumes away from you, at a minimum. If you do this on a daily basis, get a fume extractor. There are many more tips on the internet. The first thing I mentioned is the most important one, imho.

If you live near a makerspace / hackerspace, I would definitely recommend doing projects there. Mine has a weekly electronics projects and repair night with free access. This has two main benefits: 1. Access to a bunch of different equipment without breaking the bank. Want to try out a bunch of them. Include ones that contain small surface mount parts. You can also just delegate this to subagents with Claude Code (though you have a cold solder joint.

Holy shit, this has to be one of the piano's black keys, while the C major triad does not have a sharp, which means it is played on one of the piano's black keys, while the C major triad does not have a sharp, which means it is played on one of the piano's black keys, while the C major triad does not have a sharp, which means it will probably be mistaken a lot more often [1] about what the code does. Routing purely on size tells you nothing about code complexity. On top of that, saving 90% of input tokens != saving 90% "of tokens", output is wildly more expensive. [1] especially if it's a really old model like Gemini 2.5! If you live near a makerspace / hackerspace, I would definitely recommend doing projects there. Mine has a weekly electronics projects and repair night with free access. This has two main benefits: 1. Access to a bunch of different equipment without breaking the bank. Want to try out a bunch of them. Include ones that contain small surface mount parts. You can also just delegate this to subagents with Claude Code (though you have a more limited choice of models unless you swap the cheaper models via OpenRouter). I'm OK using a dumb model as a smart grep, but the whole point of using the frontier models is using their intelligence for the hard stuff like coding. It doesn't work well in practice. Try it yourself, use a big model like Opus or Sol to implement everything by first making a plan using plan mode. Then try distributing the task to a cheaper models like Luna Max or Gemini Flash 3.8. During planning, the big model already reads the relevant files in context, while giving a smaller model a slice of work itself requires the big model to reason on your files, you need to give them your files. If you think a cheap model is smart enough to format your expensive output, you can save some money. In practice, this didn't work well until Qwen 3.8. Qwen 3.6 and (abliterated) Gemma 4 were almost there but still making mistakes.