Artificial Intelligence (AI) is one of the most important emerging technologies which is transforming several human-facing sectors such as healthcare, e-commerce, social media, finance and criminal justice. As this transformation gathers pace, it is very important to ensure that the deployed AI systems are aligned with the goals and values of the humans using these systems. It is also crucial to ensure trust, safety and controlability in AI systems so that these models benefit the society and do not cause undue harm. This project aims to study the safety and value-alignment of AI models in important areas such as recommendation systems and large language models. Towards this end, the PI expects to contribute to the development of algorithms that can ensure utility-driven usage of AI models that are able to represent diverse perspectives and adhere to societal values. The proposed project aims to build India’s capacity in the emerging field of responsible and ethical AI, by understanding alignment of AI models with human preferences.