[ACL 2025] Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
-
Updated
Jun 4, 2025 - Python
[ACL 2025] Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
[ACL 2025] Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
[ACL 2025] A curated list of papers and resources based on "PlanGenLLMs: A Modern Survey of LLM Planning Capabilities"
[ACL'25 Main] SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence! | 让你的LLM更好地利用上下文文档:一个基于注意力的简单方案
[ACL 2025] SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science🌾
Embedding language models in probability space via log-likelihood vectors
[ACL 2025] Official Pytorch implementation of LADDER: Language-Driven Slice Discovery and Error Rectification in Vision Classifiers
[ACL 2025] Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information
Tandem: Riding Together with Large and Small Language Models for Efficient Reasoning (ACL 2025 Findings)
[ACL 2025 Main] Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models
Base ACL is a modular access control list API built with AdonisJS v6 that provides a robust foundation for authentication and role-based access control. The API follows clean architecture principles with clear separation of concerns and is designed to serve as a base for multiple projects.
🥷🏻 Code for our ACL 2025 Main paper: "Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates"
This is the repository of the MyMy team, ranked Top 2 in Subtask 1 and Top 2 in Subtask 2 at the SemEval 2025 Task 9: Food Hazard Detection Challenge.
(Findings of ACL 2025) TabXEval: an exhaustive, explainable rubric + two-phase framework (TabAlign → TabCompare) for table evaluation with TabXBench.
Official implementation for the ACL 2025 Main paper "Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models"
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
This repository is for the paper UAlberta at SemEval-2025 Task 2: Prompting and Ensembling for Entity-Aware Translation. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), pages 1709–1717, Vienna, Austria. Association for Computational Linguistics.
The first benchmark environment for Sensitivity Awareness (SA) in LLMs. Evaluating how language model agents handle Role-Based Access Control (RBAC), Confused Deputy vulnerabilities, and Contextual Authorization rules.
To associate your repository with the acl2025 topic, visit your repo's landing page and select "manage topics."