I pointed it at a JS and python codebase I have been working on and asked it to review and refactor, then I got it to add some features - and boy, I am not impressed.
TBH it looks the part with the running commentary it makes, but it seems to make some dumb mistakes.
Compared to claude-code it is slow, time consuming, hungry on the tokens and multiple times it told me it was done and the app could not even start.
I explicitly, and without any ambiguity, have standing orders that there is no type of silent failure that is acceptable in the current foreseen lifetime of our universe, and it still embedded a bunch of them in what was a relatively small amount of code I got it to generate.
I am just not seeing that it comes close to claude-code running Anthropic models, based on this one particular codebase I have tried it on so far.