Health Impacts of the World Trade Center Disaster-A Call to Study Those Exposed at a Young Age.
Authors: Reibman J, Trout D, Karwowski M, Wilson L, Howard J
Journal: American journal of industrial medicine
mental health
psychology
open access
Abstract
In dynamic environments animals continuously select actions to maximize a utility function. Namely, we select actions in an attempt to balance the acquisition of reward, now or in the foreseeable future, with the expenditure of resources (e.g.,energy). To enable adaptive behavior, these policies driving action selection ought to be sensitive to changes in the value of available options, our uncertainty about the state of the world, and/or our ability to alter the environment. How the brain instantiates policies—e.g., deciding what actions are worth an effort —is not fully understood. In particular, we do not fully understand how strategic, long-term planning (e.g., beyond a single trial, or offer) is updated and maintained within dynamic, closed action-perception loops. Significant progress toward answering this question has come from reinforcement learning (RL) and the development of artificial neural network agents. In this setting, we can construct various network architectures, define utility functions, and examine learned policies. This approach, when applied to neuroscience, first led to the conjecture that dopamine-driven plasticity in the striatum translates experienced action-reward associations into optimized behavioral policies. More recently, it has been suggested that a second neural RL algorithm exists, a meta-reinforcement learning system that is bootstrapped yet independent from the dopaminergic striatum. This latter meta-RL network is hypothesized to (1) be reflected in neural activity (as opposed to synaptic weights) of the pre-frontal cortex, (2) be more model-based than the dopaminergic striatum (but see), and (3) more readily allow for generalization (see for further detail). Prior work has largely supported claims casting the prefrontal cortex as implementing components of an RL algorithm. For instance, this area (or sub-divisions within) reflect(s) subjective value, decision confidence, the effort required to harvest a reward, and/or task strategy selection. These seminal studies, however, have important limitations. First, most (primate) studies artificially dissociated periods of action, perception, and reward. Second, they typically only allow for small and discrete state-spaces (e.g., two-alternative forced choice).