Independent researchers were quick to test the claims. Early results broadly match what the company said, with the usual caveat: benchmarks measure narrow skills, and real work is messier than a test set.
We will update this story as more independent testing is published. If you have used the feature and want to share what you found, write to the editorial team.
How we reported this
The announcement came with a short blog post, a technical document and a handful of demos. As usual, the demos show the best case. We read the documentation to see what the product does on an ordinary day.
- Available first to paid plans
- Usage limits apply to long tasks
- No change to default privacy settings
- API access at a separate price
What critics say
Prices did not change, but limits did. Heavy users may hit the new caps sooner, especially when they upload large files or run long tasks in the background.
We are seeing steady progress, not a leap. That is still useful for the people who rely on these tools every day.
An independent researcher
Not everyone is convinced. Some analysts say the update is mostly about catching up with rivals rather than moving the field forward. Others point to small but useful changes that the headlines missed.

What to watch next
Privacy groups asked how the new data will be stored and for how long. The company says customer data is not used for training by default on business plans. Consumer plans have a setting that you should check.
The timing is not an accident. Competitors have shipped similar updates in recent months, and every lab wants to be the default choice when a company signs a yearly contract.
At a glance
| Question | Short answer |
|---|---|
| Who is affected? | Most everyday users, gradually |
| Does it cost more? | Not for now |
| What should I do? | Try it on a low-stakes task first |
What happened
If you use these tools at work, ask your administrator whether the new feature is turned on, and what the company policy says about using it with customer data.
The practical advice is the same as always: try it on a low-stakes task first, check the output against a source you trust, and keep a human in charge of anything that matters.


