Problem 1
A)
Analysis:
For data that are not from US,those status from Mars and Venus, and those price are 0, we remove these data; some data's price has '$' dollar sign, we replace the '$' with '';
For some data, there appears "region_1": "N/A" "country": "N/A" "country": "" etc.
We can use the technique we learnt from the course to find the empty values of the fields
B)
Analysis:
x axis was choose with scalelinear at first, from the data distribution, we can see that it's more suitable to use scalelog, therefore I change method into scalelog instead. Now the data is spread more evenly into the whole image.
For the axis formatting, the original number is clear enough for readers to understand, therefore I choose the format to be .tickFormat("")
As for margin values, colors, they are choosen according to the course code. Even though we have changed the dataset we use, a good code is suitable for many datasets. That's also why refactoring code is important. Clean code makes our past work reusable and in general save time and energy.
C)
Analysis:
I choose colorscale we used in class to represent data from different stats. The advantage of jittering is to reduce overlap. The disadvantage is that the accuracy of data is influenced by this random bias number.
D)
Analysis:
The possible benefits to users from this approach is that users can see the title of the wines they are interested in conveniently.
The places in the chart where it may be ineffective or confusing is where many circles are overlapping each other.
Problem 2