AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
19066News Items
8Top Picks
124Blogs
failedLast Run

Latest AI/ML News

19066 matching items

MIT CSAIL Research 2019-06-10 16:37 UTC Score 43.0 USR-0009-20190610-research-aca-3ad70f6f Full article

MIT simulator lets users design wide range of functional soft robots

MIT simulator lets users design wide range of functional soft robots aconner Mon, 06/10/2019 - 12:37 Article June 10 '19 Adam Conner-Simons, MIT CSAIL MIT simulator lets users design wide range of functional soft robots To get robots to do things, computer scientists often use systems called physics simulators that reflect how a robot’s actions will impact the real world. These simulators don’t work particularly well, however, when it comes to soft robots made of flexible, deformable materials. This is because the underlying physical laws of deformable objects are much more complicated, requiring a lot more computational power to simulate. But in a new paper, a team from MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) has developed a new simulator made specifically for soft robots, and have shown that it can realistically simulate an eclectic mix of robotic designs, from a crawling robot to a four-legged running robot. The simulator doesn’t just efficiently evaluate robot designs, but also provides feedback on how designs can be improved. (The system’s feedback is computed based on something called “the chain rule,” and so the team has dubbed the simulator “ChainQueen”.) The team developed a high-performance GPU implementation of the simulator that they hope to eventually make open-source. “We believe this system has the potential to dramatically accelerate the development of soft robots,” says PhD student Andrew Spielberg, one of the co-authors of the…

Cross Validated 2019-06-07 08:06 UTC Score 9.0 AI-113-20190607-social-media-f2aaded8

Centering input data for Robust PCA (RPCA)?

I know that before running Principal component analys , the input data needs to be centered around its mean (subtract the mean from each keypoint) before running the algorithm. Do I need to center my data before running robust PCA ? For instance, say I have a video from a static security camera (like in the original work). Do I need to subtract from each pixel (in each frame), its mean value over the duration of the video? p.s. any other pre-processing required for RPCA?

MIT CSAIL Research 2019-05-28 17:35 UTC Score 30.0 USR-0009-20190528-research-aca-6e3ec7c0 Full article

Wireless health monitoring system shows promise in clinical trials

Wireless health monitoring system shows promise in clinical trials aconner Tue, 05/28/2019 - 13:35 Article May 30 '19 Novartis Wireless health monitoring system shows promise in clinical trials Over the past year MIT CSAIL has worked with Novartis to test a novel technology for passive, contactless monitoring of physiological signals that may be used to monitor clinical trial patients in their homes. Developed by Professor Dina Katabi and her students, the technology consists of a Wi-Fi-like device that transmits low-powered radio signals and uses machine learning algorithms to analyze their reflections and produce physiological metrics. The device can gather data on patient mobility, gait, breathing, heart rate, sleep stages, sleep apnea, and other metrics without requiring the patient to wear sensors or change their behavior in any way. Novartis and the MIT team explored the potential use of this technology in clinical trials to collect digital biomarkers, both existing and new, and potentially allow continuous, real-time monitoring of patients in their own homes. As part of the collaboration, Novartis deployed the technology in a Novartis facility, as well as in a life sciences facility with a living lab, sleep monitoring, motion and behavior monitoring. Individuals were studied for multiple days in the lab, and their motion, breathing, sleep, and behavior were measured using the technology and compared against existing standards for such measurements. Comparison to the g…

MIT CSAIL Research 2019-05-22 14:30 UTC Score 40.0 USR-0009-20190522-research-aca-9ba8b03c Full article

This robot helps you lift objects — by looking at your biceps

This robot helps you lift objects — by looking at your biceps rachelg Wed, 05/22/2019 - 10:30 Video May 22 '19 Rachel Gordon CSAIL system can mirror a user's motions and follow nonverbal commands by monitoring arm muscles. We humans are very good at collaboration. For instance, when two people work together to carry a heavy object like a table or a sofa, they tend to instinctively coordinate their motions, constantly recalibrating to make sure their hands are at the same height as the other person’s. Our natural ability to make these types of adjustments allows us to collaborate on tasks big and small. But a computer or a robot still can’t follow a human’s lead with ease. We usually either explicitly program them using machine-speak, or train them to understand our words, à la virtual assistants like Siri or Alexa. In contrast, researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) recently showed that a smoother robot-human collaboration is possible through a new system they developed, where machines help people lift objects by monitoring their muscle movements. Dubbed RoboRaise, the system involves putting electromyography (EMG) sensors on a user’s biceps and triceps to monitor muscle activity. Its algorithms then continuously detect changes to the person’s arm level, as well as discrete up-and-down hand gestures the user might make for finer motor control. The team used the system for a series of tasks involving picking up and assembling mock…

Cross Validated 2019-05-15 17:47 UTC Score 15.0 AI-113-20190515-social-media-1b81023b

Nonparametric approach for regression with a quadratic fit

I'm trying to figure out which nonparametric test I should run on my data. My data has residuals that are not normal, so I cannot run a linear regression unless I log transform it. However, log transforming my data would make it difficult to interpret a quadratic model, so I need to run a nonparametric test similar to regression. I'm trying to compare 2 models and determine which has a better fit - a linear model (y~x) or a quadratic model (y~x+x^2). Which nonparametric approach similar to regression should I use to construct each model?

Cross Validated 2019-05-09 07:14 UTC Score 13.0 AI-113-20190509-social-media-863cfe16

a measure for MAE (of a regression)

I'm running a grid search, in order to fine-tune a NN hyper parameters. the question is: the MAE values I get from the trainings are too close. since I have the statistical attributes of the target values, is there a way to somehow come up with a starting value for MAE, where the worst regression can achieve. I think the question is not clear. the value I'm looking for, is analogous to a probability for classification problems. (example: an neural network which is supposed to classify the inputs into 8 possible classes, will have accuracy of 0.125 just by random classification of inputs. so a 12.5% accuracy is the measure for such classifier) now, I have the Mean, and StdDev (and other stats if need be) of the target values of my samples. how can I calculate a measure to judge my MAE (and not by comparing different MAE values from different trainings)?

MIT CSAIL Research 2019-05-08 14:20 UTC Score 32.0 USR-0009-20190508-research-aca-4d6aaf93 Full article

CSAIL's Daskalakis wins ACM Grace Murray Hopper Award

CSAIL's Daskalakis wins ACM Grace Murray Hopper Award rachelg Wed, 05/08/2019 - 10:20 Article May 08 '19 Rachel Gordon Constantinos (“Costis”) Daskalakis, an MIT professor and CSAIL principal investigator, has won the 2018 ACM Grace Murray Hopper Award. Constantinos (“Costis”) Daskalakis, an MIT professor and CSAIL principal investigator, has won the 2018 ACM Grace Murray Hopper Award. Announced today, the prize is awarded yearly to a computer scientist on the basis of a single recent major technical or service contribution, made at or before age 35 at the time of the contribution. Daskalakis was honored for “ proving that the computational complexity of finding Nash equilibria is the same as that of finding Brouwer fixed points, a proof since extended to several other equilibrium notions.” “By challenging equilibrium theory, his work has triggered an ongoing reshaping of our understanding of strategic behavior, showing that computation must play an essential role in the foundations of game theory and economics.” His research, a fusion of computer science, economics and game theory, focuses in part on how strategic behavior complicates large-scale technological systems. To study these systems, researchers typically use equilibrium concepts, and very prominently the concept of Nash equilibrium, which occurs when every player does the best they can given other players’ choices, so no player can benefit from unilaterally changing their choice. However, Nash’s equilibrium existe…

MIT CSAIL Research 2019-04-29 17:27 UTC Score 56.0 USR-0009-20190429-research-aca-aac386e0 Full article

Giving robots a better feel for object manipulation

Giving robots a better feel for object manipulation rachelg Mon, 04/29/2019 - 13:27 Video April 29 '19 Rob Matheson Model improves a robot’s ability to mold materials into shapes and interact with liquids and solid objects. A new learning system developed by MIT researchers improves robots’ abilities to mold materials into target shapes and make predictions about interacting with solid objects and liquids. The system, known as a learning-based particle simulator, could give industrial robots a more refined touch — and it may have fun applications in personal robotics, such as modelling clay shapes or rolling sticky rice for sushi. In robotic planning, physical simulators are models that capture how different materials respond to force. Robots are “trained” using the models, to predict the outcomes of their interactions with objects, such as pushing a solid box or poking deformable clay. But traditional learning-based simulators mainly focus on rigid objects and are unable to handle fluids or softer objects. Some more accurate physics-based simulators can handle diverse materials, but rely heavily on approximation techniques that introduce errors when robots interact with objects in the real world. In a paper being presented at the International Conference on Learning Representations in May, the researchers describe a new model that learns to capture how small portions of different materials — “particles” — interact when they’re poked and prodded. The model directly learns fr…

Andrej Karpathy Blog 2019-04-25 09:00 UTC Score 46.0 USR-0115-20190425-ai-specialis-9c771960 Full article

A Recipe for Training Neural Networks

Some few weeks ago I posted a tweet on “the most common neural net mistakes”, listing a few common gotchas related to training neural nets. The tweet got quite a bit more engagement than I anticipated (including a webinar :)). Clearly, a lot of people have personally encountered the large gap between “here is how a convolutional layer works” and “our convnet achieves state of the art results”. So I thought it could be fun to brush off my dusty blog to expand my tweet to the long form that this topic deserves. However, instead of going into an enumeration of more common errors or fleshing them out, I wanted to dig a bit deeper and talk about how one can avoid making these errors altogether (or fix them very fast). The trick to doing so is to follow a certain process, which as far as I can tell is not very often documented. Let’s start with two important observations that motivate it. 1) Neural net training is a leaky abstraction It is allegedly easy to get started with training neural nets. Numerous libraries and frameworks take pride in displaying 30-line miracle snippets that solve your data problems, giving the (false) impression that this stuff is plug and play. It’s common see things like: >>> your_data = # plug your awesome dataset here >>> model = SuperCrossValidator ( SuperDuper . fit , your_data , ResNet50 , SGDOptimizer ) # conquer world here These libraries and examples activate the part of our brain that is familiar with standard software - a place where clean API…

Cross Validated 2019-04-04 13:51 UTC Score 23.0 AI-113-20190404-social-media-6558dc2a

Handling missing data in Sequence Analysis (TraMineR) within the observation window

I'm using sequence analysis. I have a question about how to deal with missing data within the observation window. The starting point of the analysis is when respondents leave secondary school (t0). I want to examine respondents' life course over a time-span of 36 months after leaving school. The dataset contains longitudinal information of respondents toward their educational histories. I arranged the data in the 'states-sequence' (STS) format. So in each month the dataset provides information on respondents' status (for example "employed" or "training"). For 58% of the sample, the data provides information over the whole observation window. So for this group, I can tell in every single month what they are doing. The sequences of the rest of the sample are shorter. Thus, the length of the sequences is not the same for all respondents. How do I handle sequences of respondents that end before month 36? What would be the way the missing values should be handled in TraMineR?

Stationary and non-stationary variables in time series - how to difference?
Cross Validated 2019-03-19 14:57 UTC Score 12.0 AI-113-20190319-social-media-9c086e1f Full article

Stationary and non-stationary variables in time series - how to difference?

I want to predict a multivariate daily time series, the target output is the volume of packages that is send and the covariates are day specific information as weather, the distance to holidays but as well lagged values of the target variable. The target output time series is not stationary, when I difference it, it is. So my intention was to just difference every variable. However, some of my covariates are already stationary, so differencing makes them non-stationary. I am not really sure if I should difference everything or nothing or just some variables, where the latter sounds not really reasonable to me. Could you please help?

Confidence interval for Population Attributable Fraction with several strata
Cross Validated 2019-03-08 18:45 UTC Score 12.0 AI-113-20190308-social-media-94892830 Full article

Confidence interval for Population Attributable Fraction with several strata

I have used aggregated data to create a table of person-years (pys) and deaths by social class, age and sex. If we consider social class to be a modifiable factor, we can calculate the number of 'expected' deaths in a situation where the low class group has the same mortality rate as the high class group. The difference is the number of attributable deaths. In the example below this is 38, and the Population Attributable Fraction for social class is 38 / 182 = 21%. +-------+-------+--------+------+--------+--------+----------+--------------+ | Class | Age | Sex | Pys | Deaths | Rate | Expected | Attributable | +-------+-------+--------+------+--------+--------+----------+--------------+ | High | Young | Male | 100 | 10 | 0.1 | 10 | 0 | | High | Young | Female | 120 | 12 | 0.1 | 12 | 0 | | High | Old | Male | 40 | 8 | 0.2 | 8 | 0 | | High | Old | Female | 80 | 12 | 0.15 | 12 | 0 | +-------+-------+--------+------+--------+--------+----------+--------------+ | Low | Young | Male | 200 | 30 | 0.15 | 20 | 10 | | Low | Young | Female | 200 | 30 | 0.15 | 20 | 10 | | Low | Old | Male | 160 | 40 | 0.25 | 32 | 8 | | Low | Old | Female | 200 | 40 | 0.2 | 30 | 10 | +-------+-------+--------+------+--------+--------+----------+--------------+ | ALL | ALL | BOTH | 1100 | 182 | 0.1655 | 144 | 38 | +-------+-------+--------+------+--------+--------+----------+--------------+ Do you know how I would calculate a confidence interval for this fraction? It seems straightforward to calculate a P…

k-fold cross validation with multiple classes
Cross Validated 2019-03-06 10:48 UTC Score 28.0 AI-113-20190306-social-media-7fa59d66 Full article

k-fold cross validation with multiple classes

I'm working on an image retrieval system (not classification). I have 5,000 images as the data set. 500 images of this dataset are the query images used for retrieval evaluation. these 500 images represent 10 different landmarks. the retrieval evaluation requires to evaluate each landmark using the average precision. and then mean average precision is measured to evaluate the 10 landmarks. However, I have a different number of query images for each landmark. some landmarks have 200 (out of 500) images as a query image and some have only 10. I'm required to divide the 500 query images into 5 folds. My question is how to perform the k-fold cross validation when the query images for each landmark varies from 10 to 200. in other words, how to deal with k-fold cross-validation in multiple classes and the sizes of the classes are different. my work is similar to the evalaution of this paper . EDIT as an example: I have 5000 images represents 10 landmarks. I have 500 query images (out of the 5000 images). The query images are as follows: landmark 1: 50 images (out of the 500). landmark 2: 10 images (out of the 500). landmark 3: 70 images : landmark 10: 200 images. I need to measure the retrieval performance for each landmark. The required number of folds is 5. Which means the 500 are supposed to be divided into 5 folds with 100 each. My question is: how to deal with the query landmarks of different sizes when distributing them across the folds?

Cross Validated 2019-02-27 11:38 UTC Score 15.0 AI-113-20190227-social-media-c1f6651a

The expected occurences of successive draws [duplicate]

We throw a coin 1 000 000 times. How many times on average will make 13 successive? Now the problem with the naive: 1 000 000/(2^13) is that once it made 13 heads the 14 head will happen with 1/2 probability , but it will count as 2 13 successive heads. 15 will happen 4xtimes less, but it will count as 3 13 successive heads. (which is far more probable than 3x1 000 000/(2^13)) if they were happening discretely. So 14 and 15 and anything above will be counted as a 13-success with value of 2,3 and so on. Monte carlo simulation also suggests something is fishy: https://stackoverflow.com/questions...nerator?noredirect=1#comment96559899_548950720134 † So what exactly is the expected number of 13 successive heads?

Cross Validated 2019-02-08 11:07 UTC Score 9.0 AI-113-20190208-social-media-b7ad7775

Proof that the addition of a baseline to the REINFORCE algorithm reduces the variance

A widely used variation of REINFORCE is to subtract a baseline value $b$ from the return $G_t$ to reduce the variance of gradient estimation, such that \begin{align} \nabla_\theta J(\theta) & \propto \sum_s d(s|\pi_\theta) \sum_a (q_\pi(s,a)-b(s)) \nabla_\theta \pi(a|s,\theta) \\ \end{align} I haven't found any proof that the baseline reduces the variance of the gradient estimation, is there one?

Cross Validated 2019-01-28 20:06 UTC Score 15.0 AI-113-20190128-social-media-fec31c0d

Importance sampling and exponential moving average

Lets say I have got a random variable $X$ with samples $x_t\sim X$ and density $p_X(x)$ and want to compute its mean via a moving average $ \mu_{t+1}=(1-c)\mu_t + c x_t$ Assume, I can not observe $X$ directly, but instead a random variable $Y$ with density $p_Y(y)$ and samples $y_t\sim Y$ . I would like to use importance sampling to compute the mean. The naive approach is with $w_t=\frac{p_X(y_t)}{p_Y(y_t)}$ $ \mu_{t+1}=\left(1- cw_t\right)\mu_t + c w_t y_t$ This works, as long as the importance weights are small. However, if $w_t> \frac 1 c$ the above formula is obviously not correct any more. Is there a way to correct for this? It should be clear that as $w_t \rightarrow \infty$ , $\mu_{t+1}\rightarrow y_t$ //edit the approach I tried to is using a sample-size estimate, but i am not sure this is correct. The initial estimate has an estimated $1/c$ samples stored. Assuming we can interpret $w_t$ as sample-size correction, we can just try to average according to how many samples we got: $ \mu_{t+1}=\left(1- \frac{w_t}{w_t + \frac 1 c -1}\right)\mu_t + \frac{w_t}{w_t+\frac 1 c-1} y_t$ for $w_t=1$ this gives the original update $\frac 1 c-1$ is the totally stored number of samples in the path after "forgetting the oldest" but i have no idea how to show that this is correct, this is just an educated guess

Cross Validated 2018-12-22 22:09 UTC Score 9.0 AI-113-20181222-social-media-ee087bc6

Maximum likelihood joint probability distribution (discrete & continuous)

I am trying to find the values $v_1$ and $v_2$ that maximizes the likelihood of some observations. I have information about $v_1$ and $v_2$ from a set of 'experiments'. In each experiment, $v_1$ and $v_2$ are corrupted with zero mean Gaussian noise, and then compared to each other, and the maximum of the two is reported. So if I have six experiments, I end up with six binary values. The standard deviation of the noise is $\sigma_c$ , equal for all experiments. I also have additional (independent) information about $v_1$ and $v_2$ . I observe two samples, $v^o_1$ and $v^{o}_2$ , sampled from a Gaussian distribution with mean equal to $v_1$ and $v_2$ (respectively) and standard deviation $\sigma_v$ (equal for the two observations). The problem is the following. I need to maximize $p(v\mid v^o, \text{experiments}, \sigma_v,\sigma_c)$ . Yet this is maximal when $v=v^o$ and $\sigma_v$ goes to zero. This occurs because the probability density goes to infinite at this point. Not sure how to deal with this. Any pointer in the right direction is much appreciated.

MIT CSAIL Research 2018-11-29 18:52 UTC Score 38.0 USR-0009-20181129-research-aca-ed14a407 Full article

Reproducing paintings that make an impression

Reproducing paintings that make an impression rachelg Thu, 11/29/2018 - 13:52 Video November 29 '18 Rachel Gordon CSAIL's new RePaint system aims to faithfully recreate your favorite paintings using deep learning and 3-D printing. The empty frames hanging inside the Isabella Stewart Gardner Museum serve as a tangible reminder of the world’s biggest unsolved art heist. While the original masterpieces may never be recovered, a team from MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) might be able to help, with a new system aimed at designing reproductions of paintings. RePaint uses a combination of 3-D printing and deep learning to authentically recreate favorite paintings — regardless of different lighting conditions or placement. RePaint could be used to remake artwork for a home, protect originals from wear and tear in museums, or even help companies create prints and postcards of historical pieces. “If you just reproduce the color of a painting as it looks in the gallery, it might look different in your home,” says Changil Kim, one of the authors on a new paper about the system, which will be presented at ACM SIGGRAPH Asia in December. “Our system works under any lighting condition, which shows a far greater color reproduction capability than almost any other previous work.” To test RePaint, the team reproduced a number of oil paintings created by an artist collaborator. The team found that RePaint was more than four times more accurate than state-of…

Cross Validated 2018-11-19 19:08 UTC Score 18.0 AI-113-20181119-social-media-25d2f345

Survival model for an epidemic -- can the observations be treated as independent?

I've been thinking about ways to tackle an epidemic modelling problem I've been working on, and I've come up against a conceptual difficulty over the way survival analysis works. Here's a really simplified version with all the extraneous details stripped out. There is a collection of $n$ individuals. At the beginning (time $t = 0$ ), 1 individual is infectious with a disease, and the other $n-1$ individuals are susceptible. As time progresses, susceptible people can get the disease through contact with infectious individuals, and pass the disease on to other people. To keep focus on the core of my question we'll use these (unrealistic) simplifying assumptions: There are no births, deaths, immigration, or emigration Individuals mix uniformly (e.g. no preferential mixing by age or sex or location, etc) No delay between contracting the disease and ability to infect others (i.e. no incubation period) Once infected, an individual remains infectious forever (no recovery) Here is how transmission works: we use a continuous-time additive hazard model. This means that we let each uninfected individual $i$ have their own hazard function $h_i(t)$ . This function tells us "given that the individual $i$ has survived until the time $t$ , what is the infinitesimal rate of failure (i.e. infection) at time $t$ ?" We define it as follows: $$h_i(t ;\lambda) = \lambda \sum_{j \neq i}^n I_j(t)$$ Where $I_j(t)$ is simply an indicator function that is 1 when individual $j$ is infectious at time $t…

Cross Validated 2018-11-14 14:00 UTC Score 9.0 AI-113-20181114-social-media-e491d14e

Expected value of quotient of Poisson distributions

Let $X$ and $Y$ be independent random variables such that $X \sim \text{Poisson}(\lambda \cdot c)$ and $Y \sim \text{Poisson}(\lambda \cdot (1-c))$ , where $c$ is a real number in $[0, 1]$ . Is there an easy way of proving that $E\left[\frac{X}{X + Y} | X + Y > 0\right] = c$ ? Numerical computations I made indicate that this is true.

Cross Validated 2018-10-08 05:03 UTC Score 12.0 AI-113-20181008-social-media-bdad0e0c

What is the difference between a Random Vector (Joint r.v.) and a Random Process?

What is the difference between a Random Vector (Joint r.v.) and a Random Process? Kindly, explain with a simple example (like toss of a coin, roll of a die, picking a card, etc.). Note. As far as I understand, random process is a collection of equally spaced and indexed (generally, by time) random variables.

Cross Validated 2018-07-17 12:08 UTC Score 12.0 AI-113-20180717-social-media-9bc1290f

Statistical significance test for comparing two canonical correlation analyses

I have a colleague who is comparing several different treatments of data via canonical correlation analysis. In other words, given some time-varying signal, $a(t)$, he extracting some vector of features $v_1(t)$. He then supposes that this is a predictor for some other vector $p(t)$. To check this he computes the [first] canonical correlation coefficient, $R_1 = \text{CCA}(v_1(t),p(t))$. And then he tries some new improved feature extractor, $v_2(t)$, and compares again $R_2 = \text{CCA}(v_2(t),p(t))$. I can find tests for the situation that $R_1$, and $R_2$ are different from zero. But what about the test that $R_1$ and $R_2$ are significantly different from each other? Asked on behalf of a colleague, but I'd also be interested in the answer.

Cross Validated 2018-05-16 11:17 UTC Score 15.0 AI-113-20180516-social-media-94c5f1a2

How does the shape of a decision boundary in relate between the original and kernel feature space?

I'm trying to get my head around the mathematics and implementation of SVM and hopefully gain some intuition into how kernels work and perhaps being able to, with a bit more confidence, define my own kernel. So far, what I understand (and I might be wrong) is that kernels map the samples in the training set into higher dimensions where the classes in the training set are linearly separable. We can then, in this new dimension use a hyperplane to separate the classes. A hyperplane is, for example, a straight line, in 2 dimensions, or a flat surface (a plane), in 3 dimensions. Questions If the original feature space is in 2 dimensions and the mapping is to 3 dimensions, then, in short, I'm wondering: When mapping a training set, generated around a quadratic function, which originally has features in the form $(X,Y)$ to $(X,Y,X^2)$, see below, doesn't seem linearly separable when plotted in a 3D space. Yet the decision boundary in the original space makes a perfect parabola and classifies samples correctly. How come? Can we and, if so, how can we visualize the decision boundary as it looks in the "extended" feature space? Can we then see how the looks in the extended feature space relate to the looks original space? Experiment In order to get some image of what this looks like I attempted the following, see github for more details: Generated 130 data points around a simple quadratic equation (2 features), Points above the quadratic curve were assigned one class points below anot…

Cross Validated 2018-05-06 19:21 UTC Score 9.0 AI-113-20180506-social-media-6612d89e

Convolution and deconvolution of random variables of different dimensions

Preliminary: Let's say we have $Y=X+Z$ ($Y$ is data, $X$ is latent variable and $Z$ is noise), where the random variables are all in $\mathbb{R}$. Then an inverse Fourier transform leads to \begin{align} f_X(x)=\frac{1}{2\pi}\int e^{-itx}\varphi_X(t)dt=\frac{1}{2\pi}\int e^{-itx}\frac{\varphi_Y(t)}{\varphi_Z(t)}dt, \end{align} where $\varphi$ are the characteristic functions. Now, suppose we know the noise distribution for $Z$ so we know the characteristic function $\varphi_Z$. If we observe data $\boldsymbol{y} = (y_1, ..., y_n)$ then we can estimate $f_Y$ and thereby estimate $\varphi_Y$ by subsitution. This gives us the estimator $$\hat{f_X}(x) = \frac{1}{2\pi}\int e^{-itx} \frac{\hat{\varphi_Y}(t)}{\varphi_Z(t)}dt = \frac{1}{2\pi}\int e^{-itx} \frac{\int e^{ity}\hat{f}_Y(y)dy}{\varphi_Z(t)}dt.$$ My question: What if $X$ is one-dimensional but the noise and observation are two-dimensional? For example, $X$ is a random variable that takes values on some deterministic line segment (in two dimensional space, but since it is on the line so essentially its one-dimensional), the error $Z$ is a two dimensional random variable so as $Y$. I understand that if we treat $X$ as a two dimensional object we still have $\varphi_X(t)=\varphi_Y(t)/\varphi_Z(t)$, but it says nothing about the fact that $f_X$ is a density on the one dimensional line segment. How should I do deconvolution in this kind of situations?

NVIDIA Blog 2018-04-12 15:27 UTC Score 29.0 AI-055-20180412-official-ai--79c2f756 Full article

Comment on What’s the Difference Between Ray Tracing and Rasterization? by Polaristar

In reply to Nutti . Yes, but this article discusses the use of ray tracing *in games.* As in, *real-time ray tracing.* We're getting to the point where software and hardware are capable of outputting ray-traced frames at 30 or 60 times a second. This is even explained near the top of the article: "Historically, though, computer hardware hasn’t been fast enough to use these techniques in real time, such as in video games. Moviemakers can take as long as they like to render a single frame, so they do it offline in render farms. Video games have only a fraction of a second. As a result, most real-time graphics rely on another technique, rasterization." This is why it's a pretty historical event.

Cross Validated 2018-02-17 17:15 UTC Score 12.0 AI-113-20180217-social-media-97aeae64

Need advice on change point (step) detection

I have a time series with lots of steps/jumps (data file here ). A plot is given below. I would like to subtract an appropriate value for each of these square wave features to bring them back down to the baseline of the signal (i.e. remove the jumps so I get a smoothly varying signal). A median filter works really well for removing a small number of outliers in a row, but in this case I probably need a different approach since the square wave jumps can have different durations as seen. A common method I've seen for doing this is to compute first differences of adjacent samples, and look for large differences to detect jumps. I implemented this method but the problem is it often fails, since the one tunable parameter for the method is a threshold value $t$ which the first differences must cross in order to detect a jump: $$ | x_{i+1} - x_i | > t $$ As can be seen in the plot below, the jumps I have are often different sizes, so a constant threshold value isn't the best approach. In particular, in some cases there is an interesting signal where adjacent samples can change by large values without being a jump! I have highlighted such a region in red. Below is a zoomed in view of the red box area. You can see there is a square wave jump followed by an interesting signal. The red arrows depict a place where adjacent samples from an interesting signal have a larger distance between them than some of the jumps in the signal. Therefore a constant threshold method with finite differe…

Cross Validated 2018-02-14 21:10 UTC Score 26.0 AI-113-20180214-social-media-0f366996

Constructing a unified path analysis model from several datasets each with different combinations of variables

I have four observational experiments (data sets) that I wish to combine and summarize in a single path analysis model. Each experiment is 3-dimensional but the observables and therefore dimensions/variables are different in each data set. One variable X.1 is common to all datasets. All other variables are present in exactly two experiments. Here are the experiments and associated variables: Experiment 1: {X.1, X.2, Y} Experiment 2: {X.1, X.2, Z} Experiment 3: {X.1, X.3, Y} Experiment 4: {X.1, X.3, Z} For the sake of my study, Y and Z are considered to be the dependent (response) variables in a set of multiple linear regression models (MLRs). For each data set I minimized the Bayesian Information Criterion (BIC) to select the best MLR model. These BIC-minimizing models (one for each experiment, same order as above) are as follows: Y ~ X.1 + X.2 Z ~ X.1 Y ~ X.1 + X.3 Z ~ X.3 I am looking for a way to create a single path analysis model with two outputs, Y and Z , that unifies all four experiments. Is there a smart way to get a single path analysis model out of this? To make such a path model I believe I must first either 1) combine these separate MLR models OR 2) combine the datasets then make a single MLR model. I'm not sure where to start for either of these methods. Thank you for reading and offering advice. BTW, there is a little bit of multicollinearity present in these MLR models, but the Variance Inflation Factor is less than 5 for all predictor variables.

Cross Validated 2018-02-05 21:35 UTC Score 9.0 AI-113-20180205-social-media-c84aedf5

Random Forest for regression--binary response

Kind of a broad question here. But is it okay/possible in R to use a random forest for regression when the response variable is a binary outcome? Essentially what I'm looking for is a probability of something happening. Below is the code I've been using it run it and the warning message. Or am I better off using this in classification mode. My end goal is to predict the probability of something happening, not necessarily predict what "classification" it will be, 1 or 0. m1RF I apologize if this question is in the wrong place.

Andrej Karpathy Blog 2018-01-20 11:00 UTC Score 25.0 USR-0115-20180120-ai-specialis-628db382 Full article

(started posting on Medium instead)

The current state of this blog (with the last post 2 years ago) makes it look like I’ve disappeared. I’ve certainly become less active on blogs since I’ve joined Tesla, but whenever I do get a chance to post something I have recently been defaulting to doing it on Medium because it is much faster and easier. I still plan to come back here for longer posts if I get any time, but I’ll default to Medium for everything short-medium in length. TLDR Have a look at my Medium blog .

MIT CSAIL Research 2017-10-17 17:58 UTC Score 35.0 USR-0009-20171017-research-aca-09c88a8a Full article

Using artificial intelligence to improve early breast cancer detection

Using artificial intelligence to improve early breast cancer detection Anonymous (not verified) Tue, 10/17/2017 - 13:58 Article October 17 '17 Adam Conner-Simons Every year 40,000 women die from breast cancer in the U.S. alone. When cancers are found early, they can often be cured. Mammograms are the best test available, but they’re still imperfect and often result in false positive results that can lead to unnecessary biopsies and surgeries. Every year 40,000 women die from breast cancer in the U.S. alone. When cancers are found early, they can often be cured. Mammograms are the best test available, but they’re still imperfect and often result in false positive results that can lead to unnecessary biopsies and surgeries. One common cause of false positives are so-called “high-risk” lesions that appear suspicious on mammograms and have abnormal cells when tested by needle biopsy. In this case, the patient typically undergoes surgery to have the lesion removed; however, the lesions turn out to be benign at surgery 90 percent of the time. This means that every year thousands of women go through painful, expensive, scar-inducing surgeries that weren’t even necessary . How, then, can unnecessary surgeries be eliminated while still maintaining the important role of mammography in cancer detection? Researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL), Massachusetts General Hospital, and Harvard Medical School believe that the answer is to turn to ar…

MIT CSAIL Research 2017-05-01 14:25 UTC Score 40.0 USR-0009-20170501-research-aca-4331151e Full article

Detecting walking speed with wireless signals

Detecting walking speed with wireless signals Anonymous (not verified) Mon, 05/01/2017 - 10:25 Video May 01 '17 Adam Conner-Simons | Rachel Gordon We’ve long known that blood pressure, breathing, body temperature and pulse provide an important window into the complexities of human health. But a growing body of research suggests that another vital sign – how fast you walk – could be a better predictor of health issues like cognitive decline, falls, and even certain cardiac or pulmonary diseases. We’ve long known that blood pressure, breathing, body temperature and pulse provide an important window into the complexities of human health. But a growing body of research suggests that another vital sign – how fast you walk – could be a better predictor of health issues like cognitive decline, falls, and even certain cardiac or pulmonary diseases. Unfortunately, it’s hard to accurately monitor walking speed in a way that’s both continuous and unobtrusive. Professor Dina Katabi’s group at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) has been working on the problem, and believes that the answer is to go wireless. In a new paper, the team presents “WiGait,” a device that can measure the walking speed of multiple people with 95 to 99 percent accuracy using wireless signals. The size of a small painting, the device can be placed on the wall of a person’s house and its signals emit roughly one-hundredth the amount of radiation of a standard cellphone. It builds o…

Disrupt Africa 2017-01-19 10:19 UTC Score 20.0 USR-0197-20170119-regional-new-5cfc4104 Full article

Comment on Dubai fintech accelerator to assist African startups by Dejene Mulugeta

I have new innovative technological solution for drinking water invisible leakage and contamination control system for each house holds and others tap water users. The technology solution is new in the global water sector. I need global financial support and partnership for my project. Many tanks!

Alignment Newsletter 2017-01-08 22:42 UTC Score 25.0 USR-0153-20170108-ai-specialis-8ee5ca74 Full article

Teaching from Simple Abstractions

(You need to know programming to understand this post. If you know what linked lists are, that’s enough to get the general point, but more knowledge would be more helpful.) Within the Programming Languages community, there’s a subcommunity that thinks a lot about education, especially for introductory courses. Two main approaches are SICP approach and […]

Alignment Newsletter 2016-12-20 21:16 UTC Score 25.0 USR-0153-20161220-ai-specialis-9217b3d4 Full article

Thoughts on the “Meta Trap”

Cross-posted to the EA Forum. Thanks to Ajeya Cotra and Jeff Kaufman for feedback on a draft of this post. Any remaining errors are my own. Last year, Peter Hurford wrote a post titled ‘EA risks falling into a “meta trap”. But we can avoid it.’ Ben Todd wrote a followup that clarified a few […]

Disrupt Africa 2016-12-18 17:01 UTC Score 20.0 USR-0197-20161218-regional-new-ba469015 Full article

Comment on 20 startups to pitch to investors at Angel Fair Africa by Dejene Mulugeta

I am working on new innovative technological solution for drinking water saveing system from invisible leakage and contamination control for each house holds and other users. The technological solution is new in the filed and also the world. Now at a time I am finishing prototype and finishing pilot test by domestic level with best result. So how can you advice me to join the global market and to get global partnership and any other support for my project? Pls.email dejenemulugeta50@gmail.com Telephone +251 922844504 Many tanks!

Disrupt Africa 2016-12-11 16:42 UTC Score 25.0 USR-0197-20161211-regional-new-5afec214 Full article

Comment on Kenyan recruitment startup Fuzu raises $1.88m by Dejene Mulugeta

Here I have new innovative technological solution for drinking water saveing systems from invisible leakage and contamination control system for each house holds and other tap water users. the technological solution is new in the field and also the world. so I need funding and partners or any other support for my project. Please email to me dejenemulugeta50@gmail.com Telephone +251 922844504 Many Tanks!

Alignment Newsletter 2016-10-20 23:08 UTC Score 25.0 USR-0153-20161020-ai-specialis-7d970b73 Full article

Almost Vegan

About a year and a half ago, I decided to stop consuming animal products, because of the intense suffering on factory farms. (Why focus on this problem among myriads of others? Jacy from Animal Charity Evaluators explains.) For the most part, I actually found being vegan easier than I thought it would be, though it still wasn’t […]

Andrej Karpathy Blog 2016-09-07 11:00 UTC Score 36.0 USR-0115-20160907-ai-specialis-85602144 Full article

A Survival Guide to a PhD

This guide is patterned after my “Doing well in your courses” , a post I wrote a long time ago on some of the tips/tricks I’ve developed during my undergrad. I’ve received nice comments about that guide, so in the same spirit, now that my PhD has come to an end I wanted to compile a similar retrospective document in hopes that it might be helpful to some. Unlike the undergraduate guide, this one was much more difficult to write because there is significantly more variation in how one can traverse the PhD experience. Therefore, many things are likely contentious and a good fraction will be specific to what I’m familiar with (Computer Science / Machine Learning / Computer Vision research). But disclaimers are boring, lets get to it! Preliminaries First, should you want to get a PhD? I was in a fortunate position of knowing since young age that I really wanted a PhD. Unfortunately it wasn’t for any very well-thought-through considerations: First, I really liked school and learning things and I wanted to learn as much as possible, and second, I really wanted to be like Gordon Freeman from the game Half-Life (who has a PhD from MIT in theoretical physics). I loved that game. But what if you’re more sensible in making your life’s decisions? Should you want to do a PhD? There’s a very nice Quora thread and in the summary of considerations that follows I’ll borrow/restate several from Justin/Ben/others there. I’ll assume that the second option you are considering is joining a medium…

Andrej Karpathy Blog 2016-05-31 11:00 UTC Score 59.0 USR-0115-20160531-ai-specialis-fd04d0db Full article

Deep Reinforcement Learning: Pong from Pixels

--> This is a long overdue blog post on Reinforcement Learning (RL). RL is hot! You may have noticed that computers can now automatically learn to play ATARI games (from raw game pixels!), they are beating world champions at Go , simulated quadrupeds are learning to run and leap , and robots are learning how to perform complex manipulation tasks that defy explicit programming. It turns out that all of these advances fall under the umbrella of RL research. I also became interested in RL myself over the last ~year: I worked through Richard Sutton’s book , read through David Silver’s course , watched John Schulmann’s lectures , wrote an RL library in Javascript , over the summer interned at DeepMind working in the DeepRL group, and most recently pitched in a little with the design/development of OpenAI Gym , a new RL benchmarking toolkit. So I’ve certainly been on this funwagon for at least a year but until now I haven’t gotten around to writing up a short post on why RL is a big deal, what it’s about, how it all developed and where it might be going. Examples of RL in the wild. From left to right : Deep Q Learning network playing ATARI, AlphaGo, Berkeley robot stacking Legos, physically-simulated quadruped leaping over terrain. It’s interesting to reflect on the nature of recent progress in RL. I broadly like to think about four separate factors that hold back AI: Compute (the obvious one: Moore’s Law, GPUs, ASICs), Data (in a nice form, not just out there somewhere on the int…

Oxford Machine Learning Research Group 2016-01-11 17:55 UTC Score 39.0 USR-0027-20160111-research-aca-96a14953 Full article

aisp_new

Autonomous Intelligent Systems This project intertwines Bayesian inference, model-predictive control, distributed information networks, human-in-the-loop and multi-agent systems. The project focuses on the principled handling of uncertainty for distributed modelling in complex environments which are highly dynamic, communication poor, observation costly and time-sensitive. We aim to develop robust, stable, computationally practical and principled approaches which naturally accommodate these rea…

Disrupt Africa 2015-12-22 13:06 UTC Score 17.0 USR-0197-20151222-regional-new-b9f45806 Full article

Comment on Banks face extinction if they don’t find ways to work with fintech startups by Dejene Mulugeta

here I have new innovative technological solution for drinking water invisible leaking and contamination control system for each house hold and other users. the technological solution is new on the field and the world so how can i get support. tanks. best regards. Dejene Mulugeta