← Back to Research Papers

Ocean acidification, more than warming or heatwaves, constrains shoaling behaviour in a range-extending fish through habitat simplification.

Authors: Mitchell A, Connell SD, Hart ME, Harvey BP, Agostini S, Spatafora D, Izumiyama M, Booth DJ, Ravasi T, Nagelkerken I
Journal: The Journal of animal ecology
mental health psychology open access

Abstract

Driven by the popularization of services such as TikTok, Instagram Reels, and YouTube Shorts, short form videos have quickly established themselves as a cornerstone of mobile platform content consumption. Unlike long-form videos, short-form videos focus on brief story telling, fast visual stimulation, and high interaction with users, giving rise to extremely dynamic consumption styles and rapidly changing user interests. These attributes create significant challenges for recommendation systems because the system must model multimodal content features while continuously updating user preference representations. The conventional recommendation systems, such as collaborative filtering, matrix factorization, etc. have been well studied and proved to be successful in static or mildly dynamic scenarios. But these approaches rely on the past user item interaction and rarely take into account the vast semantic information of short video content. Compared to traditional media, short videos naturally involve multiple heterogeneous modalities including visual appearance, audio signals and textual descriptions, all of which provide complementary information in terms of content semantics and potential user interests. Disregarding these modalities or handling them in a homogenized way may severely reduce the expressiveness and performance of recommendation models. Promising developments in multimodal recommender systems have been made to incorporate external content features to enhance recommendation quality. However, a number of existing methods either concentrate on a single leading modality (e.g., visual features) or apply early fusion or late fusion methods, which cannot explicitly capture the interactions between modalities. Besides, the majority of multimodal recommendation approaches focus on optimizing short term goals (i.e. immediate clicks or views), and they do not explicitly be built to model long term user engagement or preference drift, which is an essential feature for short-form video platform as user’s interest can volatilize very fast in the flow of time.