The subfields are defined by what they consume
The cleanest way to keep these straight is to ask what kind of input each one deals with. One branch is about text and speech, the meaning behind words and the ability to respond in kind. Another is about pixels, identifying what appears in an image or a video feed. A third, broader one is simply about learning patterns from examples rather than being explicitly programmed.
Once you sort them by input rather than by application, the questions become much easier, because most real products combine several and the question is asking which one is doing the specific job described.