AI attitudes: social-media labeling pipeline
Weibo/Twitter processing with explicit label validity, provider provenance, and resumable intermediate results.
Updated
Contents
The measurement problem
Comparing attitudes across platforms requires knowing how each observation became a label and how those labels became a daily measure. This repository processes previously extracted Weibo and Twitter data using a shared model and prompt, then cleans, aggregates, and exports figures and their underlying values. It is not a crawler.
Using the same labeling configuration controls one source of variation. It does not by itself prove cross-language equivalence, eliminate model bias, or make social-media samples representative of populations.
Processing
The client implementation pins the upstream provider and disables fallback. Each result includes provider and token metadata. Valid numeric labels and the substantive label cannot tell are distinguished from invalid or failed results.
Resumption skips observations with valid completed labels and leaves failures eligible for another attempt. Buffered labels are written as separate Parquet parts using a temporary file and rename, reducing the risk of a partially written result file. The test suite includes malformed responses, retry behavior, failed observations, and skipping completed IDs.
The README also documents daily aggregation and unsmoothed CSV values alongside plots, making it possible to inspect what smoothing changes.
A documented contribution
The September 27 storage commit introduces partitioned Parquet output, shared reading of those parts, and launch scripts. GitHub attributes it to houx15 and records an AI coauthor. The patch supports a concrete claim of documented pipeline work; it does not establish sole authorship of the research project.
Status and limits
This portfolio review did not execute the pipeline, inspect private datasets, or validate its labels against human annotations. Reproducible processing supports an audit; it does not replace a measurement-validity study or a sampling argument.