Model benchmarks measure open-ended reasoning. Support chat mostly isn't that. The answers your customers need are in your refund policy, your setup steps, your compatibility table - facts no model knows regardless of size.
That's why grounding the AI in your own documents usually improves answers more than upgrading the model does. A smaller local model with good documentation will routinely beat a frontier model guessing at your product.
Where the bigger model genuinely wins: messy, ambiguous questions; customers writing in a second language; and following long multi-step conversations without losing the thread. If that describes your traffic, weight quality more heavily.