The verified answer is C. To break text into smaller units for processing. AWS defines tokenization as the process of breaking input text into smaller units called tokens. These tokens can be words, subwords, or characters, and tokenization is used as a preprocessing step for natural language processing tasks. This exactly matches option C.
Tokenization is fundamental because machine learning and generative AI systems do not process raw human language in the same way humans read sentences. Text must first be split into units the model can represent, count, embed, and process mathematically. For example, a sentence might be split into words, subword pieces, punctuation, or characters depending on the tokenizer and model architecture. These tokens can then be converted into numerical representations for downstream tasks such as classification, summarization, question answering, translation, or text generation.
Option A is incorrect because encryption protects data confidentiality by transforming readable text into unreadable ciphertext. Tokenization in NLP is not encryption. The original purpose is linguistic and computational processing, not data protection.
Option B is incorrect because compression reduces file size. Tokenization may reduce text into units, but its purpose is not storage compression. It prepares text for analysis by NLP systems.
Option D is incorrect because translation is a downstream NLP task that converts text from one language to another. Tokenization may be used before translation, but it is not itself translation.
The wording “break text into smaller units” is the defining phrase. Therefore, the correct answer is C.