Alessandro for microsoft
We focus on advancing the state of the art in deep learning and beyond to answer fundamental questions on intelligence. We pursue research along the following directions:. Our team strives to build machines that learn from and understand the world.
We are a leader in using deep learning to solve complex language-understanding problems and in training machines to model reasoning and decision-making capabilities. Follow us:. However, since on-chip storage resources are precious and scarce Multilanguage Parallel Programming more.
The availability of high-speed user-level networking interfaces has shifted the performance bottleneck to the distributed-object infrastructures. In this paper, we describe an approach to building high-performance, commercial distributed In this paper, we describe an approach to building high-performance, commercial distributed object systems over system area networks SANs with user-level networking.
Describes an approach to build commercial high-performance distributed object systems over system area networks SANs with user-level networking.
We give a detailed functional and performance analysis of DCOM and apply optimizations at several layers to take full advantage of. Parallel processing with agora more. Parallel Processing. Speech Recognition and Distributed System. Data transfer , Partial Reconfiguration , and Upper Bound.
Energy reduction with run-time partial reconfiguration abstract only more. Reconfigurable custom floating-point instructions abstract only more. Multimedia and communication algorithms from the embedded system domain often make extensive use of floating-point arithmetic.
Due to the complexity and expense of the floating-point hardware, these algorithms are usually converted to Due to the complexity and expense of the floating-point hardware, these algorithms are usually converted to fixed point operations, or implemented using floating-point emulation in software. This study presents the design and implementation of custom floating-point units, leveraging the partial reconfiguration feature of state-of-the-art FPGAs.
The custom floating-point units can be dynamically configured, loaded, and executed when needed by software applications. The system is binary compliant with the conventional MIPS architecture and the IEEE standard, and supports most of the floating-point operations and relevant functionalities. Furthermore, we investigate various customization strategies and construct a set of optimized functional modules to meet different application demands or requirements.
Using LINPACK as a floating-point intensive example, we replace a sequence of 25 instructions with a custom unit, and demonstrate an overall 80x application speedup. Combining multicore and reconfigurable instruction set extensions more. The shift to multi-core processors presents a number of opportunities and challenges to different research fields, including the field of FPGA applications.
This paper investigates the advantages of combining multi-core processors and This paper investigates the advantages of combining multi-core processors and reconfigurable instruction set extensions. Both our analysis and the experimental results show that these two approaches exploit different levels of parallelism.
Using a case study on the Floyd-Warshall. Embedded and Case Study. Random decision tree classification is used in a variety of applications, from speech recognition to Web search engines. Decision trees are used in the Microsoft Kinect vision pipeline to recognize human body parts and gestures for a more Decision trees are used in the Microsoft Kinect vision pipeline to recognize human body parts and gestures for a more natural computer-user interface.
Tree-based classification can be taxing, both in terms of computational load and memory bandwidth. This makes highly-optimized hardware implementations. Mach: a foundation for open systems operating systems more. Introduction Operating systems have He then joined the faculty at Carnegie-Mellon University and worked first on a distributed speech recognition system Agora , and then on the Mach operating system, the OS that Apple uses today.
In , he joined Microsoft Research, betting that consumer electronics was where the next wave of system research would be.
In , he was part of the group who built the first interactive TV system codename Tiger and its Rialto operating system. Sandro was part of the team that created that watch, in MSR. Sandro has worked on networking, and was part of the team that shipped Winsock Direct in