HeartBench banner
inclusionAI inclusionAI

HeartBench

Communication community

Description

HeartBench is an evaluation benchmark for the psychological and social sciences field, designed to transcend traditional knowledge and reasoning assessments. It focuses on measuring large language models' (LLMs) anthropomorphic capabilities in human-computer interactions, covering dimensions such as personality, emotion, social skills, and ethics.

Installation

This entry records only its repository, not the path inside it, so there is no exact command to give. Open the source below and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

README

HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs

[![License: Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-green.svg)](https://opensource.org/licenses/Apache-2.0) [![Technical Report](https://img.shields.io/badge/Technical%20Report-arXiv-b31b1b.svg)](https://arxiv.org/abs/2512.21849)


🎯 Introduction

HeartBench is an evaluation benchmark for the psychological and social sciences field, designed to transcend traditional knowledge and reasoning assessments. It focuses on measuring large language models' (LLMs) anthropomorphic capabilities in human-computer interactions, covering dimensions such as personality, emotion, social skills, and ethics.

  • Evaluation Samples: 296 multi-turn dialogues
  • Scoring Criteria (Rubric): 2,818 items
  • Scenarios: 33 scenarios (e.g., personal growth, family relationships, workplace psychology)
  • Evaluation Dimensions: 5 anthropomorphism capability categories and 15 specific anthropomorphic abilities (e.g. curiosity, warmth, emotional understanding) Learn more in our research paper.

💡 Key Features

  1. Real-World Alignment: Our dataset is built from anonymized and rewritten dialogues between real users and counselors, covering high-frequency scenarios like family relationships, personal growth, and workplace psychology. We move beyond simple fact-based Q&A by employing multi-turn dialogue evaluation. The focus is on assessing a model's ability to understand complex emotions and respond to social contexts within long conversations and their subtext, rather than its capacity for simple mimicry.
  2. Fine-Grained, Science-Based Evaluation: We have developed the "AI Human-like Capability Framework," a sophisticated evaluation system rooted in established psychological theories. This framework assesses models across 5 core capabilities and 15 fine-grained subcategories, including person