Healthcare conversational AI fails in ways generic LLM testing cannot see. What the 2026 research shows, and what medical benchmarks can and cannot tell you.