Velg hva du ønkser å vise prosjekt for.
To be creative, we need to produce something which is new, meaningful and has some sort of value. Computers are able to support humans in creative processes, but to also themselves be creative or to assess if an idea or a product is creative. A master thesis project on computational creativity can investigate any creative field matching the interests and backgrounds of the student or students (language, design, music, art, mathematics, computer programming, etc.), and concentrate on one or several aspects of computational creativity, such as the production, understanding or evaluation of creativity, or on computer systems that support human creativity.
[ Vis hele beskrivelsen ]
Read also: Writing a Master's Thesis in Computational Creativity
[ Skjul beskrivelse ]
It can be vital both for security and for mental health reasons to identify users at risk in social media. That can be users being at risk of being radicalised (as extremists, school shooters and other types of terrorists, etc.) or of being targeted by predators (grooming) or extremist organisations, as well as users showing signs of mental health issues (depression, suicide, eating disorders, etc.). This thesis work would experiment with applying various machine learners, in particular Transformer-based techniques and Large Language Models, to social media text. To fully utilise deep learning, a substantial amount of data will need to be gathered, either from previous research or specifically for the task.
Read also: Writing a Master's Thesis in Language Technology
The way we write texts give a lot of information about the background personalities of the authors: their age, gender, native language (if writing in a foreign language), if they're human or bots, and possibly their actual identity. This type of information can be used to, e.g., give fair indications of user profiles, to deduce if a text (or a part of it) has been plagiarised, or to uncover social media software misuse. The thesis could thus focus on tasks such as author profiling (what can we say about the author, e.g., their gender, age, if they're a human or a bot), author identification (did a specific author write this text?) and/or plagiarism detection (did somebody else than the author claiming the text actually write all or part of the text?), looking at textual data from, e.g., social media sites, chat rooms or parliamentary debates, and apply machine learners such as Transformer technologies and Large Language Models to the texts in order to draw conclusions about who the author behind a text is.
The thesis project would explore either emotion recognition or automated generation of artworks (music, images, videos or texts) tailored to some emotion - or a combination of those two themes. In either case the work needs to explore (general) emotion taxonomies, data (music, art, poetry, etc.) with and without emotional annotations, and techniques for emotion classification in the chosen artform(s), tentatively utilising Transformer-based (Large Language Model) technology for the task of analysing and/or generating the texts, images or music scores.
Figurative language is used when the intended meaning of a statement isn't necessarily the one shown on the surface, that is, when the language intentionally conveys secondary or extended meanings, such as sarcasm, irony and metaphor. Such intentional ambiguity is also a key part of many jokes. Understanding and generating figurative language create significant challenges for language models, as direct approaches based on words and their lexical semantics often are inadequate in the face of indirect meanings. The project could thus focus on one specific type of figurative language (e.g, sarcasm or humour), and either investigate models that could interpret such figurative language or that could generate it.
Compared to other instruments such as the piano, the field of automatic processing and generation of guitar music is relatively underdeveloped. This is mainly due to the lack of large, high-quality datasets. The main challenges this project aims to tackle are thus the lack of data and the exploration of Transformer models utilised for automatic guitar tablature transcription and/or for the generation of guitar solos (e.g., in blues). This entails exploring brand-new datasets such as GAPS and SynthTab, and addressing the overfitting to the GuitarSet dataset that is very prevalent in the field, as it is one of the few datasets with a sizeable amount of richly annotated guitar music recordings.
Automatic guitar tablature transcription entails extracting guitar-specific music annotations from pieces of audio recordings of guitar music, while guitar music generation would entail building networks that are able to generate solos having significant variations from the training data and that are capable of capturing long-term dependencies in musical data.
To be creative, we need to produce something which is new, meaningful and has some sort of value. Generative AI models such as Transformer-based Large Language Models are able to support humans in creative processes, but to also itself be creative or to assess if an idea or a product is creative. A computational creativity project can investigate any creative field matching the interests and backgrounds of the student or students (language, design, music, art, mathematics, computer programming, etc.), and concentrate on one or several aspects of computational creativity, such as the production, understanding or evaluation of creativity, or on computer systems that support human creativity.
In particular, the project can investigate the transitions between different creative artforms, e.g., generating music or images based on textual input (as in Stable Diffusion models), generating music based on images or text, generating text based on music or images, or generating videos.
Properly identifying hate speech is a pressing issue for social media sites as well as for smaller companies, clubs, and organisations that allow for user-generated content. Many such sites currently use slow, manual moderation, which mean that abusive posts will be left online for too long without appropriate action being taken or that content will be published with delay (which might be unacceptable to the users, e.g., in online chat rooms).
The project would look into previous efforts to identify hate speech and cyber bullying, as well as available flame-annotated datasets from chat rooms, online games, Wikipedia, X/Twitter, etc., and investigate ways to identify such language, using various machine learning methods such as Transformer-based techniques and Large Language Models.
Pro-eating disorder groups (pro-ED) are social media sub-cultures that encourage disordered and dangerous eating behaviours, e.g., Pro-Ana (pro-anorexia), Pro-Mia (pro-bulimia) and Thinspro (Thinspiration, a combination of “thin” and “inspiration”). Automatic detection of users sharing, supporting or following pro-ED content can provide information for understanding and preventing eating disorders, as well as for social media moderation. Data on some such users on X/Twitter have already been annotated, but to fully apply machine learning algorithms such as Transformer-based Large Language Models to the problem, more data would tentatively need to be gathered or synthetic data generated.
Large language models form the basis of almost all currently topical AI research, making it vital to identify whether the models are bias based towards a certain demographic, based on gender, ethical or social background, sexual identity, religion, age, and so on. This has triggered intense research on fair representation in language models, aiming both at building and using unbiased training and evaluation datasets, and at changing the actual learning algorithms themselves. This project could investigate methods to define, identify and quantify a certain bias, as well as develop dibiasing methods, and possibly address under which circumstances a bias in an LLM even could be desirable.
Recent decades have seen a significant surge in attacks committed by lone wolf perpetrators, that is, individuals unaffiliated with any organised group and thus inherently difficult to identify beforehand. Most recently, companies behind various Large Language Models have been criticised for the systems supporting or even encouraging lone wolf terrorism. However, investigations after events such as school shooting have shown that many attackers produce indicative online posts or handwritten manifestos before the attack. This raises the possibility that extracting indicators from the texts of, e.g., previous school shooters could aid in identifying warning signs of a potential future school shooting before it takes place.
Several models can be used to find out how users’ social media networks, behaviour and language are related to their ethical practices and personalities. Such models include Schwartz’ values and ethics model and Goldberg's Big 5 model that defines personality traits such as openness, conscientiousness, extraversion, agreeableness and neuroticism. The thesis project would investigate applying such models to social media text and highlight how the user personalities are reflected by the social networks that they participate in and develop. In order to do so, a system would need to be built, tentatively based on training / fine-tuning a Transformer-based Large Language Model while utilising social media training data.
Languages change rapidly over time and language users adapt to different situations and setting. Studying how language and communicative processes evolves is a highly multi-disciplinary task involving machine and language learning as well as biological and cultural evolution. The aim of this project will be to use methods such as evolutionary algorithms or reinforcement learning to investigate the main dynamics in language evolution.
When trying to understand the origins of languages, we can compensate for the lack of empirical evidence by utilising evolutionary computational methods and/or Large Language Models to investigate how language may have evolved over time, e.g., by creating "language games" to simulate communication between agents in a social setting. In general, simulations on language evolution tend to have relatively small and fixed population sizes, something this study could aim to change.
Social media posts often express sentiment (positive or negative emotions) towards a product, person, political party, etc. The project is aimed at training / fine-tuning machine learners such as Large Language Models for the automatic classification of sentiment in texts on social media or for tracking changes in a population's opinions and attitudes. The main goal would be to determine whether a user (or group of users) has expressed a positive or negative sentiment and possibly the degree to which this sentiment has been communicated, tentatively addressing issues involving use of negation and/or sarcasm.