Web probing is an adaptation of a cognitive interview administered in a Web survey setting to understand respondents’ thought processes in formulating their answers and/or to improve survey questions. It often uses open-ended questions, which require coding the responses, an essential step for transforming qualitative data into analyzable formats. This is labor-intensive and difficult, and more so for probing on subjective and evaluative assessments, like self-rated health (SRH). Recent advancements in large language models (LLMs) offer a promising avenue for automation of this step. However, in multilingual and cross-cultural contexts, where linguistic nuances further complicate coding, their effectiveness remains insufficiently examined. This study examined LLM autocoding of open-text responses to a Web probe on the SRH question from surveys conducted across five countries (the U.S., Mexico, Germany, Great Britain, and Spain) and in four languages (English, Spanish, German, and Korean). We evaluated the performance of five major open-source LLMs by comparing their autocodes against human codes for SRH attributes (e.g., health behaviors, illness) and tone of the coded attributes (i.e., favorable, unfavorable, or neutral in relation to health). We also compared the performance of the five LLMs against traditional machine learning models (MLMs). The best-performing LLM resulted in a micro-F1 score of 0.76 for attribute classification and 0.77 for tone classification, outperforming MLMs without requiring task-specific training. Its performance was also consistent across English, Spanish, and German responses but lower for Korean, a language generally considered low-resourced for LLMs. Overall, our results demonstrate the promise of LLMs in reducing burdens for coding complex text data from SRH probing in multilingual survey research, which likely applies to a broad range of open-ended coding tasks in survey research.