Authors
Junko Hayashi, Kazuhiro Ito, Masae Manabe, Yasushi Watanabe, Masataka Nakayama, Yukiko Uchida, Shoko Wakamiya, Eiji Aramaki
Published in
JMIR formative research. Volume 10. Pages e99992. Aug 25, 2026. Epub Aug 25, 2026.
Abstract
Workday happiness is associated with workplace performance and burnout, but frequent questionnaire-based assessment is burdensome in real-world workplace settings. Free-text daily reports may provide a lower-burden way to monitor day-to-day changes in workday happiness.
This study aimed to examine whether daily diary text can be used to estimate longitudinal within-person changes in workday happiness and explore text-related factors associated with model performance, such as average sentence length and lexical diversity.
We collected free-text daily reports and self-reported workday happiness scores from employees in 2 Japanese companies. Company A provided training data from 92 participants over 2 months (1725 reports), and company B provided test data from 11 participants over 6 months (652 reports). We used 2 text-based approaches: a bidirectional encoder representations from transformers (BERT)-based regression model trained on company A data and a locally deployed Japanese large language model (Llama) used in a zero-shot setting. Model performance was evaluated for each participant using the Pearson correlation coefficient between self-reported and estimated workday happiness scores, with r=0.40 used as a pragmatic feasibility benchmark. Mean absolute error (MAE) and root mean squared error (RMSE) were also calculated on the original workday happiness scale from 0 to 10.
A total of 81.8% (9/11) of the participants met or exceeded the feasibility benchmark of r=0.40. Positive correlations were observed for 81.8% (9/11) of the participants; for 9.1% (1/11) of the participants, the correlation coefficient could not be calculated because the self-reported score remained constant, and 9.1% (1/11) showed a small negative correlation. Participant-level correlations ranged from -0.06 to 0.65 for the BERT model and from -0.05 to 0.81 for the local large language model. Error-based metrics also varied across participants: BERT MAE ranged from 1.04 to 2.78, and RMSE ranged from 1.25 to 3.50, whereas Llama MAE ranged from 0.76 to 4.39, and RMSE ranged from 1.06 to 4.64.
Our study suggests that daily report text supports low-burden, longitudinal estimation of individual-level workday happiness. However, performance varied across participants, and further work is needed to improve generalizability, reduce attrition, address possible measurement bias, and clarify appropriate workplace use.
PMID:
42641110
Bibliographic data and abstract were imported from PubMed on 26 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 14
- Comments 0