Call for Doctoral Students and Posdoctoral Researchers

AaltoLogo
TampereLogo

We are seeking new PhD students and postdocs to our computer vision and machine learning research teams in Helsinki and Tampere in Finland.

The positions will offer excellent opportunities to work in a team of professionals responsible for developing cutting edge computer vision and machine learning technology, and allows you to learn many aspects from fundamental research problems to concrete applications.

Applicant

We welcome applications with any research focus related to computer vision and machine learning. In addition, we are especially seeking candidates with the following research focus:

Ideal candidate has a strong programming and mathematics background. Experience in (or strong will to learn) programming with Python or C++/Java are considered as advantages.

Team and research

The doctoral candidates will be supervised by Professor Esa Rahtu (Tampere University) and Associate Professor Juho Kannala (Aalto University). We work broadly in the fields of computer vision and machine learning. We are pursuing research problems in geometric computer vision (including topics such as visual SLAM, visual-inertial odometry, and 3D scene reconstruction), in semantic computer vision (including topics such as image-based localization, object detection and recognition, and deep learning), and large-scale multi modal learning (including audio-visual sound source separation, visually guided sound generation, and audio-visual synchronization). More information of our research is available on our web pages linked below. (You can follow the links by clicking the images.)

Esa
Esa Rahtu
Juho
Juho Kannala

Visual-inertial odometry

We have developed a visual-inertial odometry method based on an information fusion framework employing low-cost IMU sensors and monocular or stereo camera. Our approach utilizes strong coupling between inertial and visual data sources which leads to robustness against occlusion and feature-poor environments. The video explains our method presented in this paper.

Image based localisation and SLAM

We have investigated machine learning based approaches for visual localization and SLAM systems. For instance, we have developed a CNN-based scene coordinate regression method for image-based localization. The new model can be trained without careful initialization, and the system achieves accurate results. Another example is a method for scalable and fully 3D magnetic field SLAM using local anomalies in the magnetic field as a source of position information.

Visually guided audio generation

The generation of visually relevant, high-quality sounds is a longstanding challenge of deep learning. Solving this challenge would allow sound designers to spend less time searching large foley databases for the sound that is relevant to a specific video scene. We approach the visually guided sound generation by shrinking a training dataset of audio spectrograms to a set of representative vectors aka. a codebook. Similar to word tokens in language modeling, these codebook vectors can be used by a transformer to sample a representation that can be easily decoded into a spectrogram and subsequently to audio stream. The video explains our approach presented in this paper

Audio-visual synchronisation

Audio-visual synchronisation is the task of determining the temporal offset between the audio (sound) and visual (image) streams in a video. In recent literature, this task has been explored by exploiting strong correlations between the audio and visual streams, e.g. in human speech and playing instruments, to provide a training signal for deep neural networks. Rather than focusing on a specialised domain, such as human speech, we explore AV synchronisation for videos of general thematic content, e.g. daily videos and live sports. For more information, take a look at our paper

sparse_selector_teaser.png
Audio-visual synchronisation requires a model to relate changes in the visual and audio streams. Prior work focused primarily on the synchronisation of talking head videos (left). In contrast, open-domain videos often have a small visual indication, i.e. sparse in space (right). Moreover, cues may be intermittent and scattered, i.e. sparse across time, e.g. a lion only roars once during a video clip.

Host institutions

Aalto University and Tampere University, are the leading universities in engineering and technology in Finland.

The Computer Science Department at Aalto provides world-class research and education in modern computer science to foster future science, engineering and society. The work combines fundamental research with innovative applications. The department is routinely ranked among the top 10 CS departments in Europe and in the top 100 globally.

The Signal Processing at Tampere University has 170 members of which 30-40% are of foreign origin. The department has held the prestiguous status of a Center of Excellence in Research (CoE) elected by the Finnish Academy of Sciences. Core areas of research include image, video and audio signal processing and analysis as well as machine learning related topics.

Compensation

The starting salary of a PhD student is ca. 2400 EUR per month and it will increase during the studies depending on the progress (up to 3100 EUR per month). The salary for a postdoctoral researcher starts typically from 3500 EUR per month, and increases based on experience.

In addition to the salary, the contract includes occupational healthcare benefits, and Finland has a comprehensive social security system. The positions are located at Aalto University and Tampere University.

How to apply

The applications are submitted to the application portals of the corresponding universities. We highly recommend you to apply through both portals, even if you have a clear preference on which university you prefer, this can be decided later in any case.

The links to the application portals:

Questions?

If you have any questions regarding the positions or the applications, please contact Prof. Juho Kannala (juho.kannala@aalto.fi) or Prof. Esa Rahtu (esa.rahtu@tuni.fi).